Fig.1

Concept

DiT

DiT stands for Diffusion Transformer: a diffusion model that replaces the usual U-Net denoiser with a plain Transformer operating on sequences of latent patches. If you know how a Vision Transformer chops an image into patch tokens, a DiT does the same to a noised…

The rest of “DiT” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library

DiT, explained · Fig. 1