Concept
Compute-optimal
Compute-optimal describes the allocation of a fixed training compute budget that minimizes final model loss, by choosing the right split between model size (parameters) and training data (tokens). Think of it like tuning a fixed cash budget across two inputs, except the two…
The rest of “Compute-optimal” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→