Diffusion Transformer (DiT)
A Diffusion Transformer (DiT) is a diffusion-model backbone that uses a transformer over latent image patches instead of the convolutional U-Net commonly used by earlier latent diffusion systems. The original DiT family conditions the transformer on the diffusion timestep and class information, then predicts the signal needed by the denoising process. DiT names an architecture inside a generative pipeline, not a diffusion objective by itself.
Origin and context
Peebles and Xie introduced DiT in a paper submitted in December 2022 and later published at ICCV 2023. Their experiments scaled model depth, width, and token count and compared models by forward-pass compute and image quality. The public Diffusers implementation preserves the original class-conditioned image pipeline, while later systems developed related but not identical transformer backbones.
Why it matters
DiT made transformer scaling techniques available to diffusion-based image generation and provided an alternative to a U-Net backbone. Independent work on Stable Diffusion 3 used a multimodal diffusion transformer with separate image and text streams, showing how the broader design could be adapted to text-to-image systems. That descendant is evidence of influence, but MMDiT should not be presented as the unchanged original architecture.
Example
In the original pipeline, a variational autoencoder converts an image into spatial latents. Those latents become patches processed by a transformer conditioned on a timestep and class label; a scheduler repeatedly applies its predictions to denoise the sample. Replacing that backbone with a transformer does not remove the VAE, scheduler, conditioning design, or iterative generation loop.
Maturity and evidence
Maturity is rated 3. DiT has a clear peer-reviewed formulation, public code, maintained independent library support, and independently developed transformer-based descendants. It remains below 4 because implementations vary in conditioning, attention layout, training objective, and modality handling; evidence for one benchmark or descendant cannot establish that every diffusion transformer shares the same scaling or quality properties.
Limits and open questions
Self-attention can be expensive as the number of latent patches grows, and a transformer backbone does not automatically reduce the number of denoising steps. Results depend on the latent representation, scheduler, conditioning, dataset, compute budget, and evaluation metric. DiT is also distinct from diffusion language models, and statements about Sora or other closed systems require their own primary evidence rather than inference from architectural resemblance.
Related terms
References
- Scalable Diffusion Models with TransformersMeta AI / arXiv · 2022-12-19 · class A
- DiTHugging Face · 2026 · class A
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisStability AI / arXiv · 2024-03-05 · class A
Last updated: 2026-09-03
This term is also covered in the Skills Atlas as diffusion models skill.
This term is also covered in the Skills Atlas as transformer architecture skill.
This term is also covered in the Skills Atlas as generative architectures skill.