Glossary · term

Diffusion Transformer (DiT)

A Diffusion Transformer (DiT) is a diffusion-model backbone that uses a transformer over latent image patches instead of the convolutional U-Net commonly used by earlier latent diffusion systems. The original DiT family conditions the transformer on the diffusion timestep and class information, then predicts the signal needed by the denoising process. DiT names an architecture inside a generative pipeline, not a diffusion objective by itself.

Training2022-12-19Wave 1 · 2023Maturity: 3/5

Origin and context

Peebles and Xie introduced DiT in a paper submitted in December 2022 and later published at ICCV 2023. Their experiments scaled model depth, width, and token count and compared models by forward-pass compute and image quality. The public Diffusers implementation preserves the original class-conditioned image pipeline, while later systems developed related but not identical transformer backbones.

Sources: s1, s2

Why it matters

DiT made transformer scaling techniques available to diffusion-based image generation and provided an alternative to a U-Net backbone. Independent work on Stable Diffusion 3 used a multimodal diffusion transformer with separate image and text streams, showing how the broader design could be adapted to text-to-image systems. That descendant is evidence of influence, but MMDiT should not be presented as the unchanged original architecture.

Sources: s1, s2, s3

Example

In the original pipeline, a variational autoencoder converts an image into spatial latents. Those latents become patches processed by a transformer conditioned on a timestep and class label; a scheduler repeatedly applies its predictions to denoise the sample. Replacing that backbone with a transformer does not remove the VAE, scheduler, conditioning design, or iterative generation loop.

Sources: s1, s2

Maturity and evidence

Maturity is rated 3. DiT has a clear peer-reviewed formulation, public code, maintained independent library support, and independently developed transformer-based descendants. It remains below 4 because implementations vary in conditioning, attention layout, training objective, and modality handling; evidence for one benchmark or descendant cannot establish that every diffusion transformer shares the same scaling or quality properties.

Sources: s1, s2, s3

Limits and open questions

Self-attention can be expensive as the number of latent patches grows, and a transformer backbone does not automatically reduce the number of denoising steps. Results depend on the latent representation, scheduler, conditioning, dataset, compute budget, and evaluation metric. DiT is also distinct from diffusion language models, and statements about Sora or other closed systems require their own primary evidence rather than inference from architectural resemblance.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-03

In the Skills Atlas

This term is also covered in the Skills Atlas as diffusion models skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as transformer architecture skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as generative architectures skill.