Fig.1

Concept

GSPO

GSPO stands for Group Sequence Policy Optimization, a reinforcement learning objective from the Qwen team for training reasoning models. Think of it as GRPO, except the importance-sampling ratio and clipping happen at the sequence level instead of per token.

The rest of “GSPO” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library