Concept
On-policy self-distillation
On-policy self-distillation is a training technique where a model learns from a teacher that is itself, just running under more favorable conditions. Classic distillation uses a separate, usually larger, pretrained teacher whose outputs are fixed.
The rest of “On-policy self-distillation” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→