Fig.1

Concept

One-sided advantage clamping

One-sided advantage clamping is a modification to policy-gradient reinforcement learning that limits how much a trajectory's advantage can push the policy in one direction, while leaving the other direction free. Recall that in methods like PPO or GRPO, each action (or…

The rest of “One-sided advantage clamping” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library