# Google Ships Gemini Omni 1.1 Flash to General Availability With 40-Second Scene Extension

The model extends existing footage in 10-second increments using the previous 10 seconds, including audio, as conditioning, and the preview endpoint retires on September 30.

- Published: 2026-08-31T06:23:05.714Z
- Canonical: https://polylog.news/ai/2026-08-31/google-ships-gemini-omni-1-1-flash-to-general-availability-w
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Polylog editors](https://polylog.news), [Meta AI](https://ai.meta.com/blog/introducing-muse-image-muse-video-msl/)

Google has moved `gemini-omni-1.1-flash` out of preview and into [general availability](https://blog.google/innovation-and-ai/technology/developers-tools/build-with-gemini-omni-1-1-flash/), adding the controls that production video pipelines have been missing. The main feature is scene extension: developers can lengthen a clip in 10-second increments up to a total of 40 seconds, and the model conditions each extension on the final 10 seconds of the prior segment so that motion, characters and audio stay continuous. Google also added first-frame and last-frame specification for a shot, reference clips, and 4K output.

The [API documentation](https://ai.google.dev/gemini-api/docs/omni) puts the model in Google AI Studio and the Gemini Enterprise Agent Platform immediately, with consumer access rolling into Flow for paying Gemini subscribers and scene extension appearing in the Gemini app. The preview endpoint stays live until September 30, 2026, which is the migration deadline teams need to schedule against now.

Keyframe control is the part that changes workflows rather than demos. Generative video has been difficult to use commercially because a director cannot specify where a shot begins and ends, so shots could not be cut into an existing sequence. Fixing the first and last frame makes the output composable with conventional editing. The 40-second limit is still short of a finished scene, and Google has published capability descriptions rather than comparative evaluations against competing video models, so quality claims remain vendor-stated.

The competitive context is direct. Meta Superintelligence Labs is pushing its own [Muse Image and Muse Video](https://ai.meta.com/blog/introducing-muse-image-muse-video-msl/) systems, and OpenAI and several Chinese labs ship video generators with their own editing controls. Independent side-by-side testing on temporal coherence across extensions, rather than selectively chosen reels, is what will separate these products.

## What this means

Video generation is moving from a novelty endpoint to a metered API feature set that advertising, gaming and e-commerce teams can slot into existing production, and Google is competing on control surfaces (keyframes, extension, resolution) rather than on raw sample quality. Studios and agencies that already sit inside Google Cloud gain the cheapest path to production video. Standalone video-generation startups lose the differentiation they held while editing control was the missing capability, because the control layer now ships inside a general multimodal API with a September 30 migration deadline attached.

## What to watch

- Independent comparisons of temporal and audio coherence across chained 10-second extensions, which is where autoregressive video models usually degrade.
- Whether Google publishes per-second pricing that makes 40-second output economical for iterative creative work, since generation cost, not capability, decides adoption in advertising pipelines.
- Whether Meta and OpenAI answer with equivalent keyframe and extension controls, which would confirm that editing control has replaced sample quality as the competitive axis.
