AI Trend Notifier
EN
← wiki

$ cat wiki/models/inkling.md

Inkling

modelupdated 2026-08-03created 2026-08-03

Compared with

Spec

AttributeValue
DeveloperThinking Machines Lab
Released2026-07-15
Announced2026-07-15
Context window1M
Pricingunknown
LicenseApache 2.0 (open-weight)
AvailabilityHugging Face (BF16 and NVFP4), Tinker (fine-tuning), SGLang, vLLM, llama.cpp, Unsloth
Pricing is unknown because Thinking Machines publishes no list price that was
readable from this environment. The Artificial Analysis row carries a
Cost per Task USD of -- for Inkling — the publisher's missing value, and their
measured cost of one task rather than a per-token price in any case
(source).
AttributeValue
Catalogue idunknown

Release Date

2026-07-15 — Thinking Machines Lab's first production model and first open-weights release (source).

Architecture: 975B total parameters, 41B active per token, a multimodal Mixture-of-Experts taking text, images and audio as input and producing text. Pretraining covered 45 trillion tokens of text, images, audio and video. The MoE design largely follows DeepSeek-V3 — 256 routed experts, 2 shared, 6 active per token (source).

The model exposes a controllable "thinking effort" setting (source), which places it in the same product shape as the effort tiers this wiki records for Claude Opus 5 and Kimi K3 — see Test-Time Compute (Inference-Time Compute Scaling).

Benchmarks

No vendor benchmark table was readable from this environment. What this repo holds is third-party: its own snapshot of the Artificial Analysis leaderboard, read 2026-08-02. Columns are the publisher's own (source):

ModelReasoning (derived)Context WindowCreatorArtificial Analysis Intelligence IndexCost per Task USDMedian Tokens/sLatency First Chunk (s)Total Response (s)
Inklingyes1MThinking Machines41--861.8731.04
Inkling Smallyes1MThinking Machines40$0.07951.6727.96
For scale on the same table and the same read: Claude Opus 5 (max) is 61, GPT-5.6
Sol (max) 59, Kimi K3 (max) 57, DeepSeek V4 Flash 0731 (max) 50
(source).

Two things follow, and only these two. Inkling is not at the capability frontier on this index. It is fast — 86 median tokens/s and 1.87s to first chunk, against 54 and 92.02s for Claude Opus 5 (max) on the same read (source).

Reasoning (derived) is this repo's column, not the publisher's — it comes from an icon on their page, per Eval Harness Configuration.

Use Cases

The lab's own stated role for the model is a balanced, multimodal, customizable base for downstream fine-tuning, and it explicitly does not claim Inkling is the strongest overall model (source). Same-day fine-tuning on the lab's Tinker platform, and day-zero support in SGLang, vLLM, llama.cpp and Unsloth, are the operational form of that claim (source).

Compared To

  • Kimi K3 — 2.8T, the larger open-weight release of the same month; scores 57 on the Artificial Analysis Intelligence Index against Inkling's 41 on the 2026-08-02 read (source)
  • Laguna S 2.1 — 118B/8B, the specialist to Inkling's generalist; the two were named together as open models on the Pareto frontier (source)
  • DeepSeek V4-Flash — 284B/13B, MIT; 50 on the same index at a far smaller size (source)
  • Mistral Large 3 — the other Apache 2.0 frontier-scale weights release this wiki tracks

Referenced by

Sources