Google AI Releases Gemini Omni 1.1 Flash Video Model
Google AI launches Gemini Omni 1.1 Flash with 40-second scene extension, first and last frame control, and 4K upscaling support.

Stock photo for illustration only, not from the actual event
- Gemini Omni 1.1 Flash introduces 40-second cumulative scene extension
- Features first and last frame control with native 4K upscaling
- Optimizes workflow with 60% faster 360p draft previews
Google AI has officially rolled out Gemini Omni 1.1 Flash, a cutting-edge video generation model built upon native multimodality to process text, image, audio, and video simultaneously. The model introduces conversational editing capabilities via the Interactions API, allowing users to modify existing footage seamlessly without needing to re-upload prior video files.
In terms of video extension capabilities, Omni 1.1 analyzes up to 10 seconds of prior context when continuing a clip, a major upgrade from predecessor models that only referenced the final second. Extensions operate in 10-second increments up to a cumulative 40 seconds, generating 3 to 10-second continuations per API call while automatically editing final frames to ensure continuous seams.
The integration of stateful video editing and extended context memory represents a pivotal shift in generative AI workflows. By eliminating the need for full re-renders on minor tweaks, creators can iterate rapidly while significantly lowering computational overhead and production expenses.

Stock photo for illustration only, not from the actual event
However, the model enforces specific operational constraints. Extensions can only append to the end of an existing clip without prepending or mid-clip insertion. Uploaded input videos must remain 10 seconds or shorter, and new spoken dialogue cannot be added when extending single-turn uploads containing speech, though multi-turn extensions via previous_interaction_id fully support conversational audio.
For resolution parameters, the response format supports 360p, 720p as the default, 1080p, and 4k resolutions. Google reports that 360p previews generate up to 60% faster and at one-third of the cost compared to 720p, establishing a draft-then-upscale production pattern where creators iterate cheaply before final rendering.
Pricing is structured at $1.50 per 1 million input tokens across text, image, video, and audio. Output pricing stands at $9.00 per 1 million text tokens and $17.50 per 1 million video tokens, with video billing calculated at 5,792 tokens per second for 720p footage, translating to approximately $0.10 per second.
Omni 1.1 Flash is currently available via the Gemini API in Google AI Studio, Gemini Enterprise Agent Platform, and Google Flow for subscribers, with prominent production users including Adobe, Figma Weave, GMI Cloud, and Runway already leveraging the system.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment