Skip to content

Goodhart's law variants

Definition

Goodhart's law variants (Manheim & Garrabrant, 2018, building on Garrabrant's earlier taxonomy) is the formalization of at least four distinct mechanisms by which optimizing a proxy metric diverges from the true goal — distinct failure modes that ambiguous appeals to "Goodhart's law" collapse together, each occurring and being mitigated differently.

Explanation

The paper's thesis is that metric overoptimization is not one failure but a family: the umbrella term hides mechanically different ways a metric-goal relationship breaks under optimization pressure, and discussion that ignores the distinctions cannot diagnose or prevent the failures. The taxonomy matters where optimization power is greatest — the paper singles out machine learning and AI alignment, since the pressure a system can direct at a proxy scales with its capability. For evaluation design the transferable rule is that every metric must be examined per-mechanism: a measure can survive one variant and be destroyed by another, so "is this Goodhartable?" is not a yes/no question but four separate ones. The staged source is the paper's abstract; the mechanism-by-mechanism treatment is in the full text it links.

Key Properties

  • At least four distinct mechanisms, not one law
  • Failure severity scales with optimization power directed at the proxy
  • Explicitly aimed at economics, policy, ML, and alignment
  • Diagnosis requires naming the variant, not citing the umbrella

Relationships

Applications

Auditing any metric that gates behaviour — eval scores, CI thresholds, leaderboards — one variant at a time; the vocabulary for recording why a measurement was or wasn't trusted.

Sources

  • https://arxiv.org/abs/1803.04585

See Also