Fig.1

Concept

On-policy distillation

On-policy distillation is knowledge distillation where the student learns from teacher feedback on trajectories the student itself generated, rather than on a fixed dataset of teacher outputs. The distinction matters for the same reason on-policy reinforcement learning…

The rest of “On-policy distillation” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library