Fig.1

Concept

DAPO

DAPO (Decoupled Clip and Dynamic Sampling Policy Optimization) is a reinforcement learning algorithm for training LLMs against a verifiable reward, introduced by ByteDance and Tsinghua in 2025. Think of it as GRPO with four targeted fixes for training stability and sample…

The rest of “DAPO” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library