ai-trend-notifier
← wiki

$ cat wiki/models/gpt-realtime-2.md

GPT-Realtime-2 (OpenAI)

modelupdated 2026-07-21created 2026-05-22

Spec

AttributeValue
DeveloperOpenAI
Released2026-05-07 (Realtime API GA)
Announced2026-05-07
Context window128,000 tokens (4× increase from GPT-Realtime-1.5's 32K)
Pricingunknown
Licenseunknown
AvailabilityRealtime API (generally available for production)
Reasoning effortAdjustable — minimal / low (default) / medium / high / xhigh
TypeRealtime voice model with GPT-5-class reasoning

Release Date

2026-05-07. Simultaneously: Realtime API exits beta → generally available for production.

Models in Suite

Three models released together:

ModelFunction
GPT-Realtime-2Voice reasoning (GPT-5-class), tool-calling, interruption handling
GPT-Realtime-TranslateLive translation: 70+ input → 13 output languages
GPT-Realtime-WhisperStreaming STT: transcribes live as speaker talks

Benchmarks

  • Big Bench Audio (audio intelligence): GPT-Realtime-2 high = +15.2% vs GPT-Realtime-1.5
  • Audio MultiChallenge (instruction following): GPT-Realtime-2 xhigh = +13.8% vs GPT-Realtime-1.5

Key Capabilities

  • Simultaneous tool calls: Can call multiple tools at once while speaking ("checking your calendar")
  • Interruption handling: Stays responsive while reasoning through complex requests
  • Adjustable latency/intelligence tradeoff: xhigh for complex requests, minimal for quick responses
  • 128K context: Supports long voice-based workflows (call center, medical documentation, extended sessions)

Use Cases

  • Production voice agents (first time GA)
  • Customer support voice bots with GPT-5 reasoning
  • Medical transcription and documentation
  • Live translation services (meetings, conferences)
  • Multilingual customer service

Compared To

ModelReasoningContextVoiceGA
GPT-Realtime-2GPT-5-class128KYesYes (2026-05-07)
GPT-Realtime-1.5GPT-4-class32KYesNo (beta)
Grok Voice APIYesYes
Gemini LiveYesYes (I/O 2026)

Significance

  • OpenAI's first voice model with GPT-5-class reasoning
  • Realtime API GA: voice agents now production-ready (previously beta/experimental only)
  • Adjustable reasoning effort for voice matches OpenAI's text model API pattern — unified UX
  • Moves realtime audio from "call-and-response" to "agentic voice interface"
  • GPT-Realtime-Translate: 70+ input languages for live translation is the broadest coverage in market

Open Questions

  • How does reasoning latency affect perceived naturalness in conversation?
  • Will reasoning effort be auto-selected in future (like o3's auto mode)?
  • Real-world accuracy of live translation for low-resource languages?

Sources

Related

Referenced by

Sources