For nearly a decade, computer vision and natural language processing were treated like two fundamentally alien disciplines. Language lived in the world of Transformers—scaling smoothly with compute and parameters. Vision lived in the world of Convolutional U-Nets—relying on hand-crafted downsampling and upsampling inductive biases. When researchers replaced the U-Net with a standard Transformer backbone in Diffusion Transformers (DiT), the generative video revolution was ignited.
The Limitations of Convolutional U-Nets
The original wave of image diffusion models (such as Stable Diffusion 1.5 and 2.1) relied on 2D Convolutional U-Net architectures. A U-Net processes images through a fixed hierarchy of downsampling convolutions followed by symmetric upsampling transpose convolutions.
While U-Nets worked well for $512 \times 512$ static images, they hit severe scaling walls when applied to high-resolution continuous video:
- Rigid Inductive Biases: Convolutions enforce local spatial priors, making it difficult for the network to model long-range temporal consistency and physical causality across hundreds of video frames.
- Poor Compute Scalability: Increasing compute FLOPS on a U-Net yielded diminishing visual returns compared to the predictable power-law scaling seen in Transformer models.
[U-Net Diffusion vs. Diffusion Transformer (DiT / Sora Architecture)]
Convolutional U-Net (Rigid Convolutions):
Latent Noise ──► [Conv Downsample] ──► [Bottleneck] ──► [Conv Upsample] ──► Denoised Image
(Diminishing returns when scaling compute to high-resolution multi-frame video!)
Diffusion Transformer (DiT):
Video Frames ──► [Space-Time Patchification (3D Cubes)] ──► Sequence of Visual Tokens
│
▼
[Standard Transformer Blocks + AdaLN]
│ (Cross-Attention with Text)
▼
[Photorealistic Denoised Video Spacetime!]
The DiT Paradigm: Spacetime Patchification
Developed by William Peebles and Saining Xie, the Diffusion Transformer treats visual latents identically to words in an LLM:
- Space-Time Patchification: A continuous video stream is compressed into a latent space via a 3D Variational Autoencoder (VAE), and sliced into a sequence of small 3D spatio-temporal visual patches (e.g. $2 \times 4 \times 4$ pixel cubes).
- Standard Transformer Blocks with Adaptive LayerNorm (AdaLN): The sequence of visual tokens passes through standard self-attention and MLP transformer layers. Diffusion timestep parameters and text prompt conditioning are injected dynamically via Adaptive Layer Normalization.
- Predictable Compute Scaling: Just like language models, DiT exhibits strict power-law scaling: increasing model parameters and training compute directly yields sharper details, realistic physics, and rock-solid temporal consistency.
Engineering Takeaway
Universal architectures always defeat specialized inductive biases at scale. By replacing specialized convolutional networks with scalable Diffusion Transformers, generative video systems achieved cinematic photorealism and physical consistency.