Fig.1

Concept

Token-gradient vectors

For a policy that generates text token by token, a token-gradient vector is the parameter gradient contributed by a single generated token: the gradient of that token's log-probability with respect to the model weights, typically scaled by its advantage. If you have worked…

The rest of “Token-gradient vectors” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library

Token-gradient vectors, explained · Fig. 1