Fig.1

Concept

Reward model

A reward model is a learned function that takes a piece of model output and returns a scalar score estimating how much a human would prefer it. Think of it as a regression model standing in for a human rater, so you can score millions of outputs without asking a person each…

The rest of “Reward model” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library

Reward model, explained · Fig. 1