Fig.1

Concept

CISPO

CISPO stands for Clipped Importance Sampling weight Policy Optimization, a policy-gradient RL variant introduced by MiniMax and used to train the proof-generation stage of MaxProof. Think of it as GRPO or PPO, except it clips the importance-sampling weight rather than…

The rest of “CISPO” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library

CISPO, explained · Fig. 1