⚡ BREAKING
Google DeepMind introduces Gemini Omni: multimodal model for video creation and editing
Google/DeepMind
DeepMind
Google DeepMind has unveiled Gemini Omni, a multimodal AI model capable of generating and editing videos from any combination of images, audio, video, and text. The first model in the family, Gemini Omni Flash, is rolling out to Google AI Plus, Pro, and Ultra subscribers via the Gemini app and Google Flow, and will be available to YouTube Shorts users at no cost.
Google DeepMind announced Gemini Omni, a multimodal model that combines Gemini's reasoning with creative capabilities to generate and edit videos from any input modality, including images, audio, video, and text. The model supports natural language video editing, maintains character consistency and physics across edits, and can generate videos grounded in Gemini's real-world knowledge of physics, history, science, and culture. It offers features like editing through conversation, transforming scenes, adding objects, and refining over multiple turns. The first model, Gemini Omni Flash, is being rolled out to Google AI Plus, Pro, and Ultra subscribers globally via the Gemini app and Google Flow, and also to YouTube Shorts and YouTube Create App users at no cost. All generated videos include SynthID digital watermarking for transparency. Developer and enterprise API access is planned for the coming weeks.
Source: Google DeepMind —
original
