$ diff inkling glm-5-2
Inkling vs GLM-5.2
Values come from Inkling and GLM-5.2, where each is cited to its source. This page states no benchmark result and ranks neither model — it puts two published specifications next to each other. Where a lab has not published a figure, the row says so rather than guessing.
Thinking MachinesInkling
- context
- 1M
- weights
- open
- $/M in
- —
- $/M out
- —
- context
- 1M
- weights
- open
- $/M in
- $1.4
- $/M out
- $4.4
What actually differs
- Context window
- Identical — both accept 1M tokens, so context length is not a reason to pick either.
- Weights
- Both publish weights — Inkling under Apache 2.0 (open-weight), GLM-5.2 under MIT (unrestricted, "no regional limits"). Check the licences rather than assuming they permit the same commercial use.
- Recency
- Inkling shipped 32 days after GLM-5.2 (2026-07-15 vs 2026-06-13).
Full spec
| Attribute | Inkling | GLM-5.2 |
|---|---|---|
| Developer | Thinking Machines | Z.ai |
| Released | 2026-07-15 | 2026-06-13 |
| Context window | 1M | 1,000,000 tokens |
| Pricing | not recorded | $1.40/M input · $4.40/M output — Z.ai's own endpoint on the OpenRouter catalogue, read 2026-08-01. Earlier revisions of this row quoted ≈$0.62–$0.77, which was not a price change: OpenRouter serves this model from 33 providers between $0.72 and $2.31 on input, and `/models` reports whichever one is routing at that instant. Four alerts in a day named $0.97, $1.19, $1.12 and $0.76 — Alibaba, SiliconFlow, an intermediate route, and StreamLake. Reseller quotes are real prices but they are not this model's price; quote the first-party figure, or quote a reseller with its name attached. |
| License | Apache 2.0 (open-weight) | MIT (unrestricted, "no regional limits") |
| Availability | Hugging Face (BF16 and NVFP4), Tinker (fine-tuning), SGLang, vLLM, llama.cpp, Unsloth | GLM Coding Plan (highest subscription tier), standalone API, open weights, NVIDIA NIM hosted |