Concept
On-policy distillation
On-policy distillation is knowledge distillation where the student learns from teacher feedback on trajectories the student itself generated, rather than on a fixed dataset of teacher outputs. The distinction matters for the same reason on-policy reinforcement learning…
The rest of “On-policy distillation” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→