Concept · Methodology Risk

Confirmation hacking: when AI evaluation confirms what you already believe.

Confirmation hacking is the tendency to design AI evaluations that confirm pre-existing beliefs rather than test them. It is a structural risk in any evaluation system. MO§ES™ mitigates it through content-free token telemetry — structural signals that operators cannot game by crafting prompts.

Definition

What is confirmation hacking in AI evaluation?

Confirmation hacking is the tendency to design or interpret AI evaluations in ways that confirm pre-existing beliefs rather than rigorously test them. It is a form of confirmation bias applied to the evaluation process itself.

In AI evaluation, confirmation hacking takes several forms:

  • Cherry-picked test cases — selecting benchmarks that favor your model or approach while ignoring ones that do not.
  • Custom evaluators that always agree — using an LLM-as-judge that is biased toward your outputs, or tuning the evaluator until it confirms your hypothesis.
  • Interpreting ambiguous results as confirmation — treating marginal improvements as proof while dismissing marginal regressions as noise.
  • Stopping evaluation when results look good — ending testing at the first positive result rather than running to statistical significance.
The Risk

Why confirmation hacking is dangerous.

Confirmation hacking produces evaluations that look rigorous but are not. The result is false confidence — believing your AI system or your operators are performing well when the evaluation was designed to produce that conclusion.

In enterprise settings, this leads to bad decisions: continuing to invest in approaches that are not working, deploying systems that are not safe, or concluding that training programs are effective when the evaluation was designed to confirm effectiveness.

How MO§ES™ Avoids It

Content-free telemetry cannot be gamed by prompt design.

MO§ES™ mitigates confirmation hacking through its measurement method. Because MO§ES™ uses content-free token telemetry — INPUT, OUTPUT, CACHE READ, CACHE WRITE counts — operators cannot craft prompts that produce favorable evaluations.

The structural signals MO§ES™ measures (Yield, Leverage, Token SNR, Construction) are emergent properties of how an operator interacts with AI across many sessions. They cannot be gamed by writing better prompts any more than a driver can game a fuel efficiency metric by driving more carefully for one trip.

Traditional evaluation

Output quality is judged by criteria that can be tuned until results look good. The evaluator and the evaluated have incentives to converge on favorable results.

MO§ES™ evaluation

Token telemetry is generated by the API, not by the operator. The operator cannot control what counts as INPUT, OUTPUT, CACHE READ, or CACHE WRITE. The metrics are structural, not interpretive.

Governance

Labels that prevent confirmation hacking.

MO§ES™ governance labels are designed to prevent confirmation hacking at the interpretation layer.

  • DERIVED — metrics are computed from telemetry, not interpreted from content. No subjective judgment enters the measurement.
  • DEVELOPMENTAL — results route workflows, not personnel actions. No incentive to produce favorable results.
  • ASSOCIATION — outcome relationships are labeled as association, not causation, unless validated through controlled experiments. Prevents over-interpreting correlational findings.
  • HYPOTHESIS — diagnostic patterns are labeled as hypotheses, not conclusions. Forces follow-up testing before action.
Related

Go deeper.

The core concept.

The governance framework that prevents misuse.

How benchmarks are selected to avoid bias.

Read the Methodology Request a Pilot