Concept
RLHF
RLHF (Reinforcement Learning from Human Feedback) is a technique for aligning a model's outputs with human preferences when those preferences are hard to write down as a loss function. It is like ordinary fine-tuning, except instead of optimizing against fixed labels, you…
The rest of “RLHF” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→