$ cat wiki/models/grok-imagine-video-1-5.md
Grok Imagine Video 1.5 (Preview)
modelupdated 2026-07-21created 2026-06-04
Spec
| Attribute | Value |
|---|---|
| Developer | xai |
| Released | 2026-06-03 |
| Announced | 2026-06-03 |
| Context window | unknown |
| Pricing | $0.08 per second of output video |
| License | unknown |
| Availability | api.x.ai (preview as of June 3, 2026) |
| Model ID | grok-imagine-video-1.5-2026-05-30 (alias: grok-imagine-video-1.5-preview) |
| Task | Image-to-video generation |
| Max resolution | 720p |
| Max clip length | 15 seconds |
| Architecture | Aurora engine — autoregressive mixture-of-experts, interleaved text/image/video/audio modalities |
Release Date
2026-06-03 (API preview)
Benchmarks / Rankings
- #1 on Artificial Analysis Video Arena Image-to-Video leaderboard (June 3, 2026)
- Faster inference than prior xAI video models
Use Cases
- Animating still images: provide a starting frame + motion prompt → cinematic video
- Camera moves, atmospheric effects, physics-consistent animation
- Social/marketing video production
- Developer integration via xAI API
Features
- Native audio sync (audio aligned to generated video)
- Faithful to source image (high source image adherence)
- Cinematic camera movement modeling
- Faster inference than grok-imagine-video-1.0
Compared To
| Model | Provider | Max Res | Audio | Leaderboard |
|---|---|---|---|---|
| Grok Imagine Video 1.5 | xAI | 720p | ✅ native | #1 Artificial Analysis (Jun 2026) |
| Gemini Omni | high | ✅ | top 3 | |
| Sora 2 | OpenAI | 1080p | limited | top 5 |
| Veo 3 | 1080p | ✅ | top 3 |
Significance
- xAI's Aurora engine extends beyond text to unified multimodal token generation — the same architecture for voice, image, and video
- Complements Grok Imagine (still image) and Grok Voice; Aurora is becoming a full multimodal stack
- First xAI model to top the Video Arena leaderboard — validates the quality of the Aurora engine's video generation against specialized competitors
Sources
Open Questions
- Is 1080p resolution planned?
- Text-to-video (not only image-to-video) via API?
- How will pricing evolve as the model exits preview?