Gemini Omni 1.1 Adds Longer, Controllable AI Video
Google’s production-ready video model adds 40-second extensions, keyframe interpolation, low-cost previews and output upscaling to 4K.
A more directed video workflow
Google has released Gemini Omni 1.1 Flash through the Gemini API, adding controls intended to move generative video from isolated clips toward iterative production workflows. The model can continue an existing scene in ten-second increments to a cumulative length of 40 seconds, using as much as ten seconds of preceding footage as context. Earlier systems in the family referenced only the final second.
Creators can provide both the opening and closing frames of a shot, with the model generating the movement between them. Google positions this for camera orbits, zooms, transitions and loops that are difficult to specify reliably with text alone. Users may also supply up to three seconds of reference video to guide motion or preserve visual context across a new scene.
The release introduces 360p draft generation, which Google says is up to 60% faster and costs one-third as much as standard 720p output. Selected results can then be produced or upscaled at 1080p or 4K. The model is available in Google AI Studio and the Gemini Enterprise Agent Platform. Scene extension is also rolling out globally to Google AI Plus, Pro and Ultra subscribers through Gemini and Flow.
Google’s model card still identifies consistency during complex edits, complicated motion and accurate text rendering as unresolved limitations. It also says the ability to alter a person’s speech in edited video remains restricted while safety work continues.
Why it matters
The consequential change is the separation of cheap iteration from expensive finishing. Video teams normally discard many attempts before approving a shot; lower-resolution drafts, explicit endpoint frames and reusable references make that process more predictable and economical. The 40-second ceiling remains far from full-scene filmmaking, but the release turns generation into a sequence that creative applications can expose as controllable production steps rather than a single prompt-and-hope operation.