Fig.1

Concept

Prefill

Prefill is the first phase of transformer inference, where the model processes the entire input prompt in a single forward pass before generating any output tokens. It contrasts with the decode phase, which produces one token at a time, each pass attending to all…

The rest of “Prefill” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library