$ diff gemini-3-6-flash inkling
Gemini 3.6 Flash vs Inkling
Values come from Gemini 3.6 Flash and Inkling, where each is cited to its source. This page states no benchmark result and ranks neither model — it puts two published specifications next to each other. Where a lab has not published a figure, the row says so rather than guessing.
Google DeepMindGemini 3.6 Flash
- context
- 1.0M
- weights
- closed
- $/M in
- $1.5
- $/M out
- $7.5
- context
- 1M
- weights
- open
- $/M in
- —
- $/M out
- —
What actually differs
- Context window
- Gemini 3.6 Flash takes 1.0M against 1M — modestly more room in a single request.
- Weights
- Inkling publishes weights (Apache 2.0 (open-weight)); Gemini 3.6 Flash is API-only. That decides self-hosting, air-gapped deployment and fine-tuning before any capability question does.
- Recency
- Gemini 3.6 Flash shipped 6 days after Inkling (2026-07-21 vs 2026-07-15).
Full spec
| Attribute | Gemini 3.6 Flash | Inkling |
|---|---|---|
| Developer | Google DeepMind | Thinking Machines |
| Released | 2026-07-21 | 2026-07-15 |
| Context window | 1,048,576 tokens (1M) | 1M |
| Pricing | $1.50/M input · $7.50/M output | not recorded |
| License | proprietary | Apache 2.0 (open-weight) |
| Availability | Gemini API, Google AI Studio, Android Studio, Vertex AI, consumer Gemini app, Google Search, GitHub Copilot | Hugging Face (BF16 and NVFP4), Tinker (fine-tuning), SGLang, vLLM, llama.cpp, Unsloth |