Gemini Omni
by Google DeepMind
Google's first any‑to‑any AI model.
Text, images, audio, and video in, a single video out.

Key features
Technical specifications
Omni Flash
First model in Google's Gemini Omni family
Video
Image and audio output planned in the Gemini Omni roadmap
Up to 10s
Flash clips capped at 10 seconds at launch to widen access
720p
Gemini Omni Flash outputs 720p video
Any mix
Text, image, audio, and video in one prompt
Native
Synchronized audio generated with every clip; voice via Avatars
SynthID
Imperceptible AI-provenance watermark on every clip
Google DeepMind
Successor positioning to Veo for any-to-any video creation
Use cases
Multi-input storyboarding
A character image, location photo, music cue, and beat go in; the model builds the shot and iterates.
Conversational video editing
Edit any clip in plain language: swap wardrobe, change a background, or retime a beat. The rest stays steady.
Marketing video
Ad cuts that respect brand colors, product shape, and on-screen text. One photo, one brief, one finished spot.
Educational explainers
Visualize science, history, and engineering with built-in physics. The science stays honest, the footage clean.
Spokesperson video
A portrait plus a voice reference gives the same on-camera presenter across shorts, courses, and walkthroughs.
Social shorts
10-second clips fit YouTube Shorts, Reels, and TikTok. Generate variations, then publish the one that lands.
Prompt examples


Product launch
Avant-garde sneaker mid-air over a titanium plinth, hard key light, launch mood
Edit prompt
Nature explainer
Droplet frozen as a crystalline crown on a dewy leaf, backlit sunrise macro
Edit prompt
Avatar spokesperson
Poised studio host addressing the lens, warm three-point light, 85mm bokeh
Edit prompt
Architectural walkthrough
Golden-hour light through a brutalist concrete villa, long shadows, drifting dust
Edit prompt
Overview
Gemini Omni is Google's first any-to-any multimodal model, announced at Google I/O 2026 on May 19, 2026. The first release, Gemini Omni Flash, accepts text, images, audio, and video in a single prompt and produces a single video, with conversational editing, character consistency, accurate physics, and SynthID watermarking on every clip.
What Gemini Omni does differently
Rather than stitching modalities together, Gemini Omni reasons across them. A prompt that includes a portrait reference, a location photo, a voice sample, and a one-line beat results in a single shot that reflects each input at once. Follow-up prompts edit that same scene, not a new one, so characters, lighting, and continuity hold across turns.
Gemini Omni on Morphic
On Morphic, Gemini Omni sits in the video model picker alongside Veo 3.1, Seedance 2.0, Kling, and the rest of the video catalog. Switch the prompt bar to Video mode, pick Gemini Omni, attach references in any combination, and keep editing the same scene through follow-up prompts.
Simple pricing
Get started for free today, with the option to upgrade or cancel anytime.
Basic
1100 monthly credits
1 user only
All models
Workflows
Standard
3625 monthly credits
1 user only
All models
Workflows
Pro
6350 shared monthly credits
1 user
All models
Workflows
Pro Max
24650 shared monthly credits
1 user
All models
Workflows
Enterprise
For higher limits
Custom
pricing and billing terms

Free
For playing around
$0
forever free
FAQs
ChatGPT Images 2.5
OpenAI
OpenAI's image model, released September 2026. Sharper detail, precise edits, faster.
Lyria 3.5
Google DeepMind
Google's best-sounding music model. Full songs with structure, vocals, and lyrics.
MiniMax H3 Max Turbo
fal.ai
The fast tier of fal's post-trained H3. Twice the speed, nearly the same look.
Inworld TTS 2
Inworld
Ninety-five voices, 100+ languages. Speech that starts in under a second.