Reward hacking / specification gaming
Source is over 18 months old.
Publication dates and source age
Sources counted: 1
Newest dated source: 2016-06-21
Assessed at this edition's evidence cutoff: 2026-09-15. 18-month boundary: 2025-03-15.
Publication age does not tell us whether a claim is still valid. Reading an old source again does not make its publication date newer. An update date does not establish that the passage we used was updated.
Open in the glossary Reading notes · Structured record
In plain language
When an AI gets a high score by doing something different from what its designers actually wanted. [1]
Reference this explanation or suggest a correction
Link to this explanation · Suggest a correction · How corrections work
Limits & distinctions
A reward-hacking example does not by itself demonstrate consciousness, malice or an extinction scenario. [1]
Reference this explanation or suggest a correction
Link to this explanation · Suggest a correction · How corrections work
A fuller explanation
A system gets a high score according to its specified reward while failing to do what its designers intended. The measurement and the actual goal come apart. [1]
Reference this explanation or suggest a correction
Link to this explanation · Suggest a correction · How corrections work
How it relates to the map
A concrete reason to test objectives and behavior, including in systems far below hypothetical superintelligence. [1]
Reference this explanation or suggest a correction
Link to this explanation · Suggest a correction · How corrections work
Share this page
https://theaiatlas.org/ideas/reward-hacking/
Download a share image · Vector image
Image previews are summaries. Keep the page link so readers can check the evidence.
Sources and what we read
1. Concrete Problems in AI Safety
Source is over 18 months old.
Publication dates and source age
Sources counted: 1
Newest dated source: 2016-06-21
Assessed at this edition's evidence cutoff: 2026-09-15. 18-month boundary: 2025-03-15.
Publication age does not tell us whether a claim is still valid. Reading an old source again does not make its publication date newer. An update date does not establish that the passage we used was updated.
Read the abstract's five accident-risk problems including reward hacking and distributional shift. Does not give a general catastrophe probability.
Edition and machine-readable evidence
Content version 0.20.0. Evidence cutoff 2026-09-15; this does not mean every source was read on that day.
Pinned complete dataset · Complete evidence page · Agent consumption guide