$ cat wiki/models/gemini-3-5-flash-lite.md
Gemini 3.5 Flash-Lite
modelupdated 2026-07-22created 2026-07-22
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | 2026-07-21 |
| Announced | 2026-07-21 |
| Context window | 1,048,576 tokens (1M) |
| Pricing | $0.30/M input · $2.50/M output |
| License | proprietary |
| Availability | Gemini API, Google AI Studio, Vertex AI |
Benchmarks
| Benchmark | Gemini 3.5 Flash-Lite | Gemini 3.1 Flash-Lite (predecessor) |
|---|---|---|
| Terminal-Bench 2.1 | 54% | 31% |
| GDM-MRCR v2 (long-context) | 72.2% | 60.1% |
| GDPval-AA v2 | 1140 | 642 |
| AA Intelligence Index | +11 pts over predecessor | baseline |
| Time per task | ~halved vs. predecessor | baseline |
| Output speed | 350.2 t/s | unknown |
Use Cases
- High-throughput, low-latency agentic pipelines (agentic search, document processing)
- Cost-sensitive applications requiring millions of calls per day
- Fast classification, routing, and summarization at scale
Key Differentiators
- Cheapest production Gemini model: $0.30/$2.50 per 1M tokens
- Speed: 350.2 t/s — significantly above the category median (~107 t/s)
- Configurable thinking levels: supports fast and deep reasoning modes
- Positioned as the high-throughput tier complementing Gemini 3.6 Flash
Related
- Google DeepMind — developer
- Gemini 3.6 Flash — higher-performance sibling, released same day
- Gemini 3.5 Flash Cyber — security-specialized sibling, released same day
- Agents (LLM Agents) — primary deployment context
Sources
- Google Blog (July 21, 2026): https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
- DeepMind Flash-Lite page: https://deepmind.google/models/gemini/flash-lite/