Fig.1

Concept

GRPO

GRPO (Group Relative Policy Optimization) is a reinforcement learning method for fine-tuning language models, introduced by DeepSeek. Think of it as PPO with the value network removed.

The rest of “GRPO” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library