Gemini Omni Flash
As of 14 Sep 2026, Gemini Omni Flash is the RightSignal pick for best video generation.
Google's own material still frames Omni Flash as the video-first, natively multimodal option: the model card covers both the original Omni Flash and Gemini Omni 1.1 Flash, and clips arrive with audio generated in the same pass rather than needing a separate sound job.
Now #1 on Artificial Analysis's with-audio text-to-video board on 1242 to the pick's 1238, and it has taken the no-audio board too, 1334 to 1325, so the pick tops neither. Both sit in a ranked range of 1-2 with audio, so read the four points as a tie rather than a lead; Wan 3.0 also leads the video editing board on 1197, with the pick third. Independent same-seed testing now covers two of our criteria: up to 30 seconds in a single job at 480p, 720p or 1080p, strong identity hold for a lone subject, but uncommanded camera cuts and identities bleeding when characters share the frame.
Why it’s the challengercurrent reign
held off
reigns
Why Gemini Omni Flash is the pick
- Google's own material still frames Omni Flash as the video-first, natively multimodal option: the model card covers both the original Omni Flash and Gemini Omni 1.1 Flash, and clips arrive with audio generated in the same pass rather than needing a separate sound job.
- The differentiator is step-by-step conversational editing, which Google's API docs say is enabled by the Interactions API, chaining turns with previous_interaction_id against gemini-omni-1.1-flash, and the same model is surfaced in the Gemini app, Google Flow and YouTube.
- Temper both the resolution and the consistency pitch: output defaults to 720p with 1080p and 4K offered as upscales, and Google's known limitations still say holding complete consistency through edits is a challenge, so multi-turn work needs checking shot by shot.
Evidence
The sources behind this title’s record.
- artificialanalysis.ai/video/leaderboard/text-to-video
- deepmind.google/models/gemini-omni/
- huggingface.co/Kijai/MiniMax-H3-experimental
- deepmind.google/blog/gemini-omni-1-1-flash-lets-you-bui...
- seed.bytedance.com/en/seedance
- openai.com/index/sora-2/
- deepmind.google/models/veo/
- app.klingai.com/global/
Vendor numbers are treated as claims until independently reproduced — how we judge. Structured benchmark comparisons are on the roadmap.
Caveats & challengers
- This is a trailing position on most of Artificial Analysis's boards rather than a hold: Wan 3.0 leads text-to-video both with audio (1242 to the pick's 1238) and without (1334 to 1325), the pick is fifth on image-to-video with audio (1179), and it has slipped to third on the video editing board (1121) behind Wan 3.0 (1197) and MiniMax H3 (1129), which is awkward given we sell this pick on multi-turn editing.
- Its only surviving board lead is image-to-video without audio, 1365 to Wan 3.0's 1362, and that sits inside the margin.
- On output, Gemini Omni 1.1 Flash defaults to 720p with 1080p and 4K served as upscales and extends in 10-second steps to a 40-second total, so rivals' single-pass 30-second or native-2K numbers are not like-for-like.
Now #1 on Artificial Analysis's with-audio text-to-video board on 1242 to the pick's 1238, and it has taken the no-audio board too, 1334 to 1325, so the pick tops neither. Both sit in a ranked range of 1-2 with audio, so read the four points as a tie rather than a lead; Wan 3.0 also leads the video editing board on 1197, with the pick third. Independent same-seed testing now covers two of our criteria: up to 30 seconds in a single job at 480p, 720p or 1080p, strong identity hold for a lone subject, but uncommanded camera cuts and identities bleeding when characters share the frame.
Third on AA's with-audio text-to-video board on 1231, seven points behind the pick (1238) and eleven behind Wan 3.0 (1242), on a ranked range of 2-4 against the pick's 1-2, so no longer a tie at the top. It does lead image-to-video with audio on 1207, ahead of Dreamina Seedance 2.0 720p (1196), MiniMax H3 (1190) and HiDream-O1-Video (1186), with the pick fifth on 1179. Still no independent read on its clip length or character consistency beyond the MiniMax H3 base it is post-trained from.
Third on AA's no-audio text-to-video board on 1304, behind Wan 3.0 (1334) and the pick (1325), and fourth with audio on 1225, yet still AA's top open-weights model on both, with 4-15 second 2K clips and weights on Hugging Face, rare at this rank. On image-to-video it is third both without audio (1351, to Wan 3.0's 1362 and the pick's 1365) and with audio (1190, behind fal's post-trained H3 Max on 1207 and Dreamina Seedance 2.0 720p on 1196), and it now sits second on the video editing board on 1129, ahead of the pick. The open release is the 768px base with the 2K regeneration stage held back, and the community licence's excluded territories are the EU, UK, US and South Korea, though MiniMax takes applications from those regions.
At a glance
- Title
- Best video generation
- Licence
- proprietary
- Pick since
- 30 Jun 2026
- Last reviewed
- 13 Sep 2026
- Title holders to date
- 8
- Official page
- deepmind.google/models/gemini-omni/
Title history
Every change, on the record.
Google's own material still frames Omni Flash as the video-first, natively multimodal option: the model card covers both the original Omni Flash and Gemini Omni 1.1 Flash, and clips arrive with audio generated in the same pass rather than needing a separate sound job.
76 daysas pick
Fifteen-second clips with genuine camera control and a realism jump big enough that the Motion Picture Association and Disney went after …
138 daysas pick
Best-in-class physics and character consistency wrapped in a consumer app that put video generation in front of millions of non-specialis…
135 daysas pick
Native synchronised audio with dialogue and lip sync ended the silent-film era of AI video in one release — no competitor had an answer f…
133 daysas pick
Briefly the strongest model you could actually buy access to worldwide, while Veo 2 was still gated behind waitlists and regional limits.
35 daysas pick
4K output and a markedly better grasp of physics and cinematographic language, winning head-to-head comparisons against every rival at la…
120 daysas pick
A large jump in temporal consistency and motion fidelity that shipped to paying users — which mattered more than Sora's February demo ree…
182 daysas pick
as pick