Fig.1

Concept

Reinforcement learning

Reinforcement learning (RL) trains a model by having it act in an environment and rewarding good outcomes, rather than showing it correct answers directly. Contrast it with supervised fine-tuning (SFT), where you hand the model labeled input/output pairs and minimize…

The rest of “Reinforcement learning” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library

Reinforcement learning, explained · Fig. 1