Fig.1

Concept

KL penalty

The KL penalty (Kullback-Leibler divergence penalty) is a regularization term that keeps your fine-tuned policy from drifting too far from a reference model during RL optimization. Think of it like L2 regularization pulling weights toward zero, except here the pull is toward…

The rest of “KL penalty” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library