Google is giving developers more control over AI-generated video with Gemini Omni 1.1 Flash, an updated model that adds longer scene extensions, keyframe-based transitions, video references and output upscaling to 4K. The focus is less on generating another impressive short clip and more on addressing the awkward parts of actually building AI video into creative workflows.
One of the biggest changes is scene extension. Gemini Omni 1.1 Flash can examine up to 10 seconds of an existing video before generating what comes next, rather than relying on only the final second of footage. Google says this additional context improves continuity and adherence to an existing scene. Videos can be extended in 10-second increments, with a cumulative maximum length of 40 seconds.
That limit still leaves Omni well short of replacing conventional editing and production tools for longer projects, but maintaining characters, environments and camera logic across successive generations has been a persistent problem for generative video. More contextual awareness could make extensions considerably more useful than simply stitching together loosely related AI clips.
Creators can also specify the first and last frames of a shot, leaving the model to generate the movement between them. That gives developers a more deliberate way to create transitions, camera moves and looping sequences instead of relying entirely on text prompts to describe where a shot should begin and end.
Google is also tackling the cost of experimentation. Omni 1.1 Flash can generate 360p previews for drafting and prototyping, which the company says can be up to 60% faster than generating at 720p while costing one-third as much. Once a sequence is ready, output can be produced at 1080p or upscaled to 4K. Those performance figures are Google’s own and will depend on how developers use the model in practice.
Another addition lets developers supply as much as three seconds of reference video alongside other multimodal inputs. This could be useful for preserving movement, character behaviour or visual context when constructing a new scene, particularly in applications where text and static reference images do not provide enough information.
The wider significance is Google’s shift toward controllability rather than raw generation alone. AI video models can already produce visually convincing footage under favourable conditions, but professional workflows demand repeatability, iteration and some ability to direct what happens between frames. Features such as low-resolution drafts, reference footage and explicit start and end frames are aimed squarely at that problem.
Gemini Omni 1.1 Flash is available through the Gemini API in Google AI Studio and Google’s enterprise agent platform. Google says the model is also available globally through Google Flow for AI Plus, Pro and Ultra subscribers, while scene extension is being offered to those subscription tiers in the Gemini app.
The technical upgrades make Omni 1.1 Flash a more flexible production tool on paper. The harder test will be whether its continuity and control remain dependable across repeated generations, rather than only in carefully selected demonstrations.

