What is DiT?
DiT (Diffusion Transformer) is a transformer-based denoising or velocity-prediction backbone that represents a spatial input, commonly a noisy image latent, as patch tokens and conditions their processing on a diffusion or flow timestep and other inputs.
Quick Facts
| Full Name | Diffusion Transformer |
|---|---|
| Created | 2023 by William Peebles and Saining Xie (Meta/UC Berkeley) |
| Specification | Official Specification |
How It Works
Follow the patchify, condition, transform, unpatchify path
For a latent grid of height H, width W, and patch size p, the original design creates (H/p) * (W/p) tokens. A linear patch embedding and positional encoding feed standard transformer blocks. Timestep and class embeddings condition the blocks; the original ablation found adaLN-Zero, whose residual modulation is initialized to produce an identity-like block, strongest within its tested setup. A final projection and unpatchify step produce the spatial denoising output.
Treat scaling claims as experiment-bound evidence
The original paper varied model depth, width, and latent patch size on class-conditional ImageNet and observed better FID as forward-pass Gflops increased in that study. This is not a universal scaling law for every dataset, objective, resolution, or hardware stack. Smaller patches increase token count; halving each patch dimension quadruples tokens and increases the number of dense attention pairs sixteen-fold before implementation details.
Distinguish later transformer backbones and objectives
MMDiT gives text and image streams separate weights while allowing joint attention, whereas the original DiT primarily used timestep and class conditioning. Other systems use cross-attention, U-shaped transformer hierarchies, token reduction, diffusion losses, flow matching, or rectified flow. Compare complete pipelines on matched data, compute, latency, memory, sampler, autoencoder, and evaluation; the shared use of transformer blocks does not make architectures equivalent.
Key Characteristics
- Uses transformer blocks as the generative model backbone rather than defining the full pipeline
- Converts spatial latents or pixels into a sequence of patch tokens
- Injects timestep and other conditions through a declared conditioning mechanism
- Trades smaller patches and longer sequences for higher compute and memory cost
- Includes later variants with cross-attention, joint modalities, or hierarchical token paths
- Requires experiment-specific quality, latency, memory, and robustness evaluation
Common Use Cases
- Testing a transformer backbone in a latent image diffusion pipeline
- Conditioning image generation on classes, text tokens, or control inputs
- Extending spatial patch sequences with temporal tokens for video research
- Studying model-size and token-count scaling under a fixed training protocol
- Comparing isotropic, multimodal, and hierarchical generative backbones
Example
Loading code...Frequently Asked Questions
Is DiT a diffusion model or only a backbone?
DiT is primarily the neural backbone that maps a noisy spatial representation and conditioning inputs to a denoising or velocity prediction. The complete generative system also chooses an autoencoder or pixel space, forward path, training target, conditioning encoder, guidance method, sampler, and data protocol.
Is a Diffusion Transformer always better than a U-Net?
No. The original result compared a controlled DiT family on class-conditional ImageNet, not every U-Net, dataset, or budget. Data scale, parameter count, token count, conditioning, training objective, implementation, and hardware affect quality and efficiency. Use matched experiments rather than architecture labels.
What does adaLN-Zero do in the original DiT?
It derives scale, shift, and residual-gating values from timestep and class conditioning and zero-initializes the residual modulation so each block begins close to an identity mapping. It performed best among the conditioning designs tested in the original paper; that result is specific to its experiment.
How is MMDiT different from the original DiT?
The original DiT study focused on class-conditioned latent patches. MMDiT gives image and text token streams modality-specific weights while allowing information exchange through joint attention. It was evaluated with a rectified-flow text-to-image system, so its architecture and training objective should not be collapsed into one generic DiT recipe.
Why does DiT patch size strongly affect compute?
For a two-dimensional latent, token count is `(H/p) * (W/p)`. Halving patch size in both dimensions creates four times as many tokens, and dense self-attention then has sixteen times as many token pairs. Actual runtime also depends on hidden size, kernels, memory traffic, sparsity, and parallelism.