You ask a video to move its subject to a forest. Then you ask it to make the violin invisible. Then you ask for an over-the-shoulder angle. Each edit preserves what came before. The violinist is still there, still playing, but now you are watching from behind them, in a forest, with no instrument in sight.
That is Google Omni. Not another video generator. A video model you can talk to.
What Omni Actually Does
Gemini Omni is Google’s first natively multimodal generation model. It takes any combination of text, image, video, and audio as input. It produces video as output. The first release, Gemini Omni Flash, is available now to Google AI Plus, Pro, and Ultra subscribers.
But the input/output specs undersell it. The real capability is conversational editing.
Every generation builds on the last. You do not start over. You do not re-prompt from scratch. You refine, redirect, restyle, and reframe in natural language, turn by turn, and every turn carries forward. No other video model does this. Kling generates. Seedance generates. Veo generates. Omni generates, then listens, then revises.
The Capabilities That Matter
Conversational Editing
This is the killer feature. Start with any video or image. Type "Move them to a forest setting." The scene relocates. The subject stays consistent. Type "Make it sunset." The lighting shifts. Type "Change to slow motion." The pacing drops. Every instruction stacks. You are editing by describing what you want changed, not by selecting clips on a timeline.
For anyone who has spent hours in Premiere moving keyframes, this is a different way of thinking about video. You describe the change. The model handles the reconstruction.
Style Transformations
Touch a mirror in a scene. Type "Make the mirror ripple like liquid." Or "Turn the person into line art." Or "Make it felted puppet style with googly eyes." Three different visual directions from the same starting point. The model understands what style means, not just what pixels to shift.
World Knowledge
This is where Omni separates from the pure-pixel models. Ask for "a claymation explainer of protein folding, accurate" and you get a stop-motion video that is scientifically correct. Ask for "a skeuomorphism stop motion explainer about how the brain hippocampus works" and the anatomy is right. The model reasons about what should happen, not just what looks plausible.
Physics
"Marble rolling on a Rube Goldberg track, continuous shot." Omni generates a physics-accurate chain reaction. Gravity, momentum, kinetic energy transfers, all correct. This is not just visual simulation. The model understands cause and effect in the physical world.
Text in Video
Alphabet items appear with correct lower-third text overlays. Word-by-word animated text syncs to timing, each word rendered in a different animated style. Text generation inside video has been the hardest problem in this space. Omni handles it.
Multi-Reference Compositing
Feed it a character image, a motion video, and an audio track. Omni composites all three into a single coherent output, with style synced to the beat. Character identity, movement, and audio all fused in one pass. No manual compositing step.
Audio-Reactive Generation
"Add harp sounds synchronized to when I touch each fern leaf. Change to bioluminescent with fireflies that react as I play." The model generates audio that matches physical interactions in the scene, and visual effects that respond to that audio. Sound and image, coupled.
Environment Restyling
Feed it a walking video. Type "Imagine the world gradually changing into retro-futuristic style as I walk." The environment transforms progressively as the person moves through it. The original footage persists. The style overlays it.
Google Flow Updates
Omni arrives inside Google Flow, Google’s creative workspace. Three additions are worth noting.
Flow Agent plans, brainstorms, batch edits, creates variations, and organizes assets. It operates across an entire project, not just one generation at a time.
Flow Tools lets you vibe-code your own creative tools without writing code. Describe what you need in natural language. A custom editor, resizer, or shader appears. Share it with others.
Flow Music with Omni directs music videos conversationally, synced to track pacing. You describe the visual treatment. The model times cuts and transitions to the audio.
What to Look Out For
Omni Flash is live now in the Gemini app, Google Flow, and YouTube Shorts. It requires a Google AI Plus subscription at $20 per month, or Pro or Ultra tiers. API access is coming in the coming weeks via Vertex AI.
All generated content carries SynthID watermarking and C2PA Content Credentials. This is Google’s answer to the provenance question: every Omni output is traceable to its source model.
The conversational editing loop is where the real shift happens. Video production has always been a render-wait-review cycle. Omni collapses that into a dialogue. You describe. It delivers. You refine. It revises. The creative feedback loop goes from minutes to seconds.
For designers and creatives, the practical takeaway is straightforward: start thinking about video the way you think about design iterations. You do not need a final concept before you begin. You can generate, react, and redirect. The model holds context across turns. Your first idea does not have to be your final one.
The Bigger Signal
Also at Google I/O: TikTok launched an MCP server for AI agents to run ad campaigns programmatically. Google, Meta, Amazon, and TikTok are all building MCP infrastructure. Every major platform is preparing for a world where AI agents, not humans, are the primary API consumers.
Omni fits this trajectory. A model that edits video conversationally is a model an agent can drive. When the interface is natural language, humans and agents use the same on-ramp.
The video model that listens is here. The question is what you ask it to make.



