Gemini Omni Flash 1.1
by Google DeepMind
Google's video model where the edit is a sentence.
Now with 4K output and 40‑second scenes.

Key features
Technical specifications
Gemini Omni
Second Omni Flash release in Google's Omni family
5 tasks
Text, image, reference, edit, and extend
3 to 10s
Per generation; extensions reach 40 seconds total
Up to 4K
360p and 720p native, 1080p and 4K upscaled
10 + 3
Ten images plus three video clips of up to 3s each
Native
Dialogue, effects, and music in the same pass
16:9, 9:16
Landscape and vertical at every resolution
SynthID
Invisible provenance mark on every generated clip
Use cases

Edits you type
Hand it a clip and say the change: swap the weather, restyle to anime, rewrite the sign. The rest of the frame stays as shot.

Talking characters
Write the line and the character says it, lip-synced, in a voice that survives every cut and every edit that follows.

Reference-true spots
A product still, a presenter photo, and a three-second clip of the move you want become one spot with the product held true.

Scenes that grow to 40s
Generate the opening beat, then extend it ten seconds at a time into a 40-second scene with the same cast, setting, and score.

Explainers with physics
Gravity, fluids, and collisions follow real-world rules, and Gemini's grounding in science and history keeps the detail right.

Titles rendered in frame
On-screen text comes back readable, from lower thirds to a storefront sign, and a follow-up edit can rewrite what it says.
Prompt examples
Overview
Gemini Omni Flash 1.1 is the second release of Google DeepMind's multimodal video model, and the first with room to grow a shot. A generation still runs 3 to 10 seconds with native audio, but 1.1 adds the controls the first release lacked: a resolution ladder from 360p drafts to 4K delivery, reference videos that genuinely steer the output, an end-frame option that turns two stills into one motion, and extension that carries a clip to 40 seconds without losing the cast or the score.
What's new in Gemini Omni Flash 1.1
The first Omni Flash rendered every clip at 720p and stopped at ten seconds. Version 1.1 keeps the fast, conversational core and removes the ceilings around it. Resolution is now a choice per generation, with 360p for cheap iteration and 1080p or 4K upscales when a take is ready to ship. Reference input grows to ten images plus three short clips, each tagged to a role in the prompt. And a finished clip is no longer the end of the line: extension reads the last stretch of footage and continues the scene, ten seconds at a time, up to four times the original length.
Conversational video editing, not re-rolling
Most video models treat a prompt as a slot machine pull. Omni Flash treats it as the first note in a session. The model holds the scene between turns, so "make the jacket red", "now extend the shot as she turns away", and "change the sign to say CLOSED" all land on the same take, with faces, wardrobe, and voices intact. Short, single-change instructions work best, and the parts of the frame you never mention are the parts that do not move.
Gemini Omni Flash 1.1 on Morphic
On Morphic, Gemini Omni sits in the video model picker alongside Veo 3.1, Seedance 2.5, Kling, and the rest of the catalog. Write the shot as a brief that includes the audio, attach reference images for anything that must stay consistent, and generate; the take lands on the Canvas, where the same brief can run through another model for a side-by-side read. Keep refining in follow-ups, one focused change at a time, and the scene keeps its context from turn to turn.
Simple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
1100 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Pro Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
ChatGPT Images 2.5
OpenAI
OpenAI's image model, released September 2026. Sharper detail, precise edits, faster.
Lyria 3.5
Google DeepMind
Google's best-sounding music model. Full songs with structure, vocals, and lyrics.
MiniMax H3 Max Turbo
fal.ai
The fast tier of fal's post-trained H3. Twice the speed, nearly the same look.
Inworld TTS 2
Inworld
Ninety-five voices, 100+ languages. Speech that starts in under a second.