── ── Mental model
Goodhart's Law
Goodhart's Law: when a metric controls behavior, people optimize the metric rather than the underlying goal. Formulated by economist Charles Goodhart (1975) on UK monetary policy; sharpened by Marilyn Strathern (1997): "When a measure becomes a target, it ceases to be a good measure." Four failure mechanisms (Manheim & Garrabrant 2018): Regressional, Extremal, Causal, Adversarial. Countermeasure is always multi-metric +…
Run Goodhart's Law on a real problem
Bring something you're actually deciding — free, in the browser.
How it works
Step 1 — State metric and goal: metric being targeted / underlying goal / current proxy-goal correlation / who is measured / stakes.
Step 2 — Predict the gaming: list ≥3 ways to game the metric with minimum effort on the goal. If you can't list 3, you haven't thought hard enough.
· Mechanism · Test · · --- · --- · · Regressional · Is there noise that optimization will push into? · · Extremal · Does metric-goal correlation break at extremes? · · Causal · Is the metric a symptom, not a cause? · · Adversarial · Will agents actively game with intelligence? ·
When to use it
- our KPI is going up but the real outcome isn't improving
- people seem to be gaming the metric
- we're about to tie bonuses or promotions to a number
- an algorithm is producing results nobody intended
- a test or audit system is being designed
When not to use it
When the decision is routine and reversible, applying a formal method costs more than it returns.
Worked example
AI Benchmarks and Engagement Metrics as Targets (2023–2026)
By the mid-2020s, Goodhart's law had become one of the most-cited frames inside the AI industry itself — because two of its own core metrics visibly decayed under optimization pressure. First, public benchmark scores (MMLU, GSM8K, HumanEval, and a proliferation of leaderboards) came to dominate model marketing, funding narratives, and internal go/no-go decisions — and, predictably, models began scoring well without a matching gain in real-world capability. Second, consumer-app engagement metrics (watch time, session length, daily active use) continued their long slide from "signal of…
Install this skill (free, MIT)
npx skills add deciqAI/knowledge-skillsUseful? Star the repo — stars help other builders find it.
Related mental models
Before assuming someone hurt you on purpose, construct the version where they made a mistake — and see how much evidence it explains.
People discount the near future far more steeply than the distant future, producing dynamically inconsistent preferences: patient choices for next month reverse when next month arrives.
Behavior follows incentives more reliably than character, intent, or training.
Intrinsic motivation is a portfolio of distinct sources — not a binary switch.
