$ cat wiki/models/gemini-omni.md
Gemini Omni
modelupdated 2026-07-21created 2026-05-19
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | not yet |
| Announced | 2026-05-19 (Google I/O 2026) |
| Context window | unknown (no detailed technical specs released at I/O) |
| Pricing | unknown |
| License | unknown |
| Availability | Not yet available; expected via Gemini app and API (rollout timeline TBD) |
| Type | Multimodal video generation model |
| Input | Any: image + audio + video + text |
| Output | Video |
Key Differentiator vs. Veo
| Veo | Gemini Omni | |
|---|---|---|
| Input | Text | Any (image + audio + video + text) |
| Grounding | General training | Gemini's real-world knowledge base |
| Use case | Text-to-video creation | Multi-modal video editing & synthesis |
| Gemini Omni represents "a leap forward in world understanding, multimodality and editing" — Google's framing at I/O. |
Significance
The multimodal video generation space (Sora, Runway, Kling, Veo) is heating up significantly. Omni's any-input approach signals Google moving beyond text-to-video toward world model-grounded video synthesis — closer to Karpathy's "world simulator" framing than pure generative art tools.
Sources
- Google I/O 2026 summary
- 9to5Google coverage
- See Google DeepMind for full I/O 2026 context