Google DeepMind introduced Gemini Omni, starting with Omni Flash for video creation. The announcement describes combining text, images, audio, and video as inputs and connects generation with Gemini’s multimodal reasoning.

Context

This resembles a creative brief expressed in several media. The remaining challenge is deciding which reference takes priority: what should stay exact, and what may change? Explicit constraints make the output easier to evaluate than an open-ended request for something attractive. Mixed inputs increase expressive control only when their roles are clear.

Sources & authors

  1. Introducing Gemini Omni
    Google DeepMind · May 19, 2026