Try the new DeepSeek V4 Pro 0813 today. Frontier intelligence at a fraction of the cost. Here
LLM

Thinking Machines Lab LogoInkling

Inkling is a multimodal MoE model (975B total, 41B active, 256k context) reasoning over text, image, and audio inputs.

Model details

Model API pricing

Per 1M tokens

See all pricing

View repository

Inkling is a general-purpose, open-weights multimodal model from Thinking Machines Lab that takes text, image, and audio inputs and generates text. Architecturally, it's a mixture-of-experts transformer with 975B total parameters (41B active), a 256k context window, and native vision and audio. It was trained for breadth rather than a single vertical.

Released under open weights, it's aimed at developers building agentic and tool-use systems, coding assistants, chatbots, and RAG applications, with support for English and other languages across multiple programming languages.

Input
1# You can use this model with any of the OpenAI clients in any language!
2# Simply change the API Key to get started
3
4from openai import OpenAI
5
6client = OpenAI(
7    api_key="YOUR_API_KEY",
8    base_url="https://inference.baseten.co/v1"
9)
10
11response = client.chat.completions.create(
12    model="thinkingmachines/inkling",
13    messages=[
14        {
15            "role": "user",
16            "content": "Implement Hello World in Python"
17        }
18    ],
19    top_p=1,
20    max_tokens=1000,
21    temperature=1,
22    presence_penalty=0,
23    frequency_penalty=0
24)
25print(response.to_json())
JSON output
1{
2    "id": "chatcmpl-xxxx",
3    "choices": [
4        {
5            "finish_reason": "stop",
6            "index": 0,
7            "message": {
8                "content": "Here is the canonical \"Hello, World!\" in Python:\n\n```python\nprint(\"Hello, World!\")\n```\n\n**How it works**\n- `print()` is Python's built-in function for sending output to the console.\n- The text inside the quotes is a string (you can use single `'` or double `\"` quotes).\n\n**Running it**\nSave the code in a file (e.g., `hello.py`) and run:\n```bash\npython hello.py\n```\n\nOutput:\n```\nHello, World!\n```",
9                "role": "assistant",
10                "reasoning_content": "The user wants a \"Hello World\" implementation in Python. This is a very basic request. I should provide the simplest, most standard version, and maybe mention a few common variations (like using print with different styles, or the classic \"Hello, World!\").\n\nThe standard is:\n```python\nprint(\"Hello, World!\")\n```\n\nOr with single quotes:\n```python\nprint('Hello, World!')\n```\n\nI can present this clearly with explanation that `print()` is the built-in function to output text to the console.\n\nSince it's so simple, I should keep it concise but clear. Maybe include how to run it briefly?\n\nLet's provide:\n1. The code\n2. Brief explanation\n3. Maybe execution instructions (optional but helpful for absolute beginners)\n\nActually, the prompt is \"Implement Hello World in Python\" - just need the code.\n\nI'll provide the canonical version with a brief explanation."
11            }
12        }
13    ],
14    "created": 1784140413,
15    "model": "inferact/inkling-nvfp4",
16    "object": "chat.completion",
17    "usage": {
18        "completion_tokens": 295,
19        "prompt_tokens": 9,
20        "total_tokens": 304,
21        "completion_tokens_details": {
22            "accepted_prediction_tokens": 0,
23            "audio_tokens": 0,
24            "reasoning_tokens": 182,
25            "rejected_prediction_tokens": 0
26        },
27        "prompt_tokens_details": {
28            "audio_tokens": 0,
29            "cached_tokens": 0
30        }
31    }
32}

🔥 Trending models