Google DeepMind introduced Gemini Omni, starting with Omni Flash for video creation. The announcement describes combining text, images, audio, and video as inputs and connects generation with Gemini’s multimodal reasoning.
Context
This resembles a creative brief expressed in several media. The remaining challenge is deciding which reference takes priority: what should stay exact, and what may change? Explicit constraints make the output easier to evaluate than an open-ended request for something attractive. Mixed inputs increase expressive control only when their roles are clear.
Sources & authors
- Introducing Gemini OmniGoogle DeepMind · May 19, 2026


