Fig.1

Concept

LLaDA

LLaDA stands for Large Language Diffusion with mAsking, a family of language models that generate text by iterative denoising instead of next-token prediction. Think of it as BERT-style masked prediction, except the mask ratio is a diffusion timestep and the model…

The rest of “LLaDA” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library