ai-trend-notifier
← wiki

$ cat wiki/models/gemini-3-5-flash.md

Gemini 3.5 Flash

modelupdated 2026-07-21created 2026-05-19

Spec

AttributeValue
DeveloperGoogle DeepMind
Released2026-05-19 (GA at Google I/O 2026)
Announced2026-05-19 (Google I/O 2026)
Context window1,048,576 input / 65,536 output tokens
Pricing$1.50 in / $9.00 out per 1M tokens ($0.15 cached)
Licenseunknown
AvailabilityGemini API, Google AI Studio, Google Antigravity, Gemini app, AI Mode in Google Search
FamilyGemini 3.5
ModalitiesText + image + audio + video → text
ThinkingDynamic thinking (on by default)
Knowledge cutoffJanuary 2026
Speed4× faster than other frontier models (output tokens/sec)

Benchmarks

BenchmarkScoreContext
Terminal-Bench 2.176.2%Agentic coding
MCP Atlas83.6%MCP/tool use
CharXiv Reasoning84.2%Chart reasoning
GPQA Diamond90.4%PhD-level science
MMMU-Pro81.2%Multimodal understanding
Outperforms Gemini 3.1 Pro across the full coding and agentic benchmark suite.

Use Cases

Primary positioning: agentic and coding tasks at Flash speed and cost. The "strongest agentic and coding model the Flash series has ever shipped."

  • Multi-step agent workflows (MCP Atlas 83.6%)
  • Scientific reasoning (GPQA Diamond 90.4%)
  • Coding tasks (Terminal-Bench 76.2%)
  • Multimodal analysis

Availability

Gemini API, Google AI Studio, Google Antigravity, Gemini app, AI Mode in Google Search.

Compared To

  • Gemini 3.1 Pro: Gemini 3.5 Flash surpasses it on coding/agentic/multimodal benchmarks at lower cost and 4× speed
  • GPT-5.5 Instant: Both target fast, cost-efficient frontier tiers. Gemini 3.5 Flash has an explicit MCP Atlas benchmark; GPT-5.5 Instant lacks published agentic benchmarks
  • Claude Opus 4.7: Anthropic flagship vs. Google's Flash tier; different price/performance tradeoffs
  • Devstral 2: Open-weight coding specialist (72.2% SWE-bench) vs. Gemini 3.5 Flash's broader multimodal + agentic scope

Sources

Referenced by

Sources