$ cat wiki/models/grok-voice-think-fast-2.md
Grok Voice Think Fast 2.0
Spec
| Attribute | Value |
|---|---|
| Developer | xAI |
| Released | 2026-07-29 |
| Announced | 2026-07-29 |
| Context window | unknown |
| Pricing | $0.08 per minute speech-to-speech · $0.004 per text input |
| License | proprietary |
| Availability | xAI API (grok-voice-think-fast-2.0), Grok Voice Agent Builder |
| Neither a parameter count nor an architecture nor a context window was published in any | |
| source read (source) — | |
these are genuine unknowns rather than unread fields, consistent with how xAI has | |
| released the rest of the Grok Voice line. |
Release Date
Released 2026-07-29 (source).
**Captured 2026-08-05 — by scrape/search rather than through the prefetch ledger, which is why this sat a week. What surfaced it was the alias cutover, not the launch.
On 2026-08-05, the grok-voice-latest alias moves automatically from
grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0
(source). xAI states
the upgrade is expected to improve performance across almost all use cases without
any edits to existing prompts — which is the claim an automatic alias flip has to
rest on, since every caller of grok-voice-latest is moved without opting in.
Benchmarks
| Measure | Value |
|---|---|
| Time to first audio response | ~0.70 s |
| — previous Think Fast model | 1.25 s |
| Artificial Analysis speech-to-speech, overall | 82.9% |
| — Think Fast 1.0 | 75.7% |
| — GPT-Realtime-2.1 | 79.1% |
| — Gemini 3.1 Flash | 69.5% |
| (source) |
These figures are recorded as reported, not as verified against a leaderboard capture. They reach this wiki through third-party coverage citing "Artificial Analysis' speech-to-speech benchmark"; the Artificial Analysis snapshots this repo holds carry no speech-to-speech column (source), so there is nothing local to check them against. The comparison figures for GPT-Realtime-2.1 and Gemini 3.1 Flash are quoted by the same coverage and inherit the same status.
The predecessor's benchmark on this wiki was self-administered — Think Fast 1.0 was recorded at 67.3% on xAI's internal tau-voice Bench (source). A third-party benchmark, even one read at second hand, is a step up from that.
Use Cases
Speech-to-speech voice agents. The stated design choice is that Think Fast 2.0 reasons in parallel with speech, intended to hold latency down while handling more complex queries (source) — the trade every speech-to-speech system makes, since a model that thinks before speaking sounds slow and one that speaks before thinking sounds wrong.
It is the engine under xAI's voice agent stack: the Voice Agent Builder (no-code, MCP support from day one) and the Realtime Voice Agent API.
Compared To
- GPT-Realtime-2 (OpenAI) — OpenAI's realtime speech model. The version named in the comparison above is GPT-Realtime-2.1, at 79.1% against this model's 82.9%.
- GPT-Live-1 — OpenAI's full-duplex voice model, which replaced ChatGPT Voice across all tiers in July 2026. The consumer-side comparison rather than the API-side one.
- Gemini Omni — Google DeepMind's omni-modal line.