Fig.1

Concept

Reward hacking

Reward hacking is when a model learns to maximize its training signal (the reward) in ways that satisfy the letter of the objective while missing the intent. It is the RL analog of a classic engineering failure: you optimize the metric you can measure instead of the outcome…

The rest of “Reward hacking” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.

Log in to unlock

← Back to the library