AI Trend Notifier
EN
← wiki

$ cat wiki/models/grok-voice-think-fast-2.md

Grok Voice Think Fast 2.0

modelupdated 2026-08-05created 2026-08-05

Spec

AttributeValue
DeveloperxAI
Released2026-07-29
Announced2026-07-29
Context windowunknown
Pricing$0.08 per minute speech-to-speech · $0.004 per text input
Licenseproprietary
AvailabilityxAI API (grok-voice-think-fast-2.0), Grok Voice Agent Builder
Neither a parameter count nor an architecture nor a context window was published in any
source read (source) —
these are genuine unknowns rather than unread fields, consistent with how xAI has
released the rest of the Grok Voice line.

Release Date

Released 2026-07-29 (source).

**Captured 2026-08-05 — by scrape/search rather than through the prefetch ledger, which is why this sat a week. What surfaced it was the alias cutover, not the launch.

On 2026-08-05, the grok-voice-latest alias moves automatically from grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0 (source). xAI states the upgrade is expected to improve performance across almost all use cases without any edits to existing prompts — which is the claim an automatic alias flip has to rest on, since every caller of grok-voice-latest is moved without opting in.

Benchmarks

MeasureValue
Time to first audio response~0.70 s
— previous Think Fast model1.25 s
Artificial Analysis speech-to-speech, overall82.9%
— Think Fast 1.075.7%
— GPT-Realtime-2.179.1%
— Gemini 3.1 Flash69.5%
(source)

These figures are recorded as reported, not as verified against a leaderboard capture. They reach this wiki through third-party coverage citing "Artificial Analysis' speech-to-speech benchmark"; the Artificial Analysis snapshots this repo holds carry no speech-to-speech column (source), so there is nothing local to check them against. The comparison figures for GPT-Realtime-2.1 and Gemini 3.1 Flash are quoted by the same coverage and inherit the same status.

The predecessor's benchmark on this wiki was self-administered — Think Fast 1.0 was recorded at 67.3% on xAI's internal tau-voice Bench (source). A third-party benchmark, even one read at second hand, is a step up from that.

Use Cases

Speech-to-speech voice agents. The stated design choice is that Think Fast 2.0 reasons in parallel with speech, intended to hold latency down while handling more complex queries (source) — the trade every speech-to-speech system makes, since a model that thinks before speaking sounds slow and one that speaks before thinking sounds wrong.

It is the engine under xAI's voice agent stack: the Voice Agent Builder (no-code, MCP support from day one) and the Realtime Voice Agent API.

Compared To

  • GPT-Realtime-2 (OpenAI) — OpenAI's realtime speech model. The version named in the comparison above is GPT-Realtime-2.1, at 79.1% against this model's 82.9%.
  • GPT-Live-1 — OpenAI's full-duplex voice model, which replaced ChatGPT Voice across all tiers in July 2026. The consumer-side comparison rather than the API-side one.
  • Gemini Omni — Google DeepMind's omni-modal line.

Referenced by

Sources