AI evaluation should learn from how we test humans.
AI evaluation focuses on models — benchmarks, safety tests, output quality. But human evaluation has evolved over a century to measure something more useful: performance in context. AI operator evaluation should borrow from that tradition.
We learned this lesson with humans a century ago.
In the early 20th century, industrial psychology discovered that capability tests do not predict job performance. A high IQ score does not mean someone will be a good employee. What predicts performance is... performance — observed behavior in the actual work context.
This is why modern human evaluation has multiple layers:
- Aptitude tests — measure capability in isolation (like model benchmarks)
- Certifications — measure demonstrated skill (like output evaluation)
- Performance reviews — measure observed behavior in real work (like operator evaluation)
No one confuses an aptitude test with a performance review. We know they measure different things. Yet in AI evaluation, we routinely confuse model benchmarks (aptitude tests) with operator performance (performance reviews). We evaluate the model and assume we have evaluated the operator.
Model benchmarks are aptitude tests.
MMLU, HumanEval, and Chatbot Arena are aptitude tests. They measure what the model can do in a controlled setting. They do not measure what your operators will do with the model in real work.
This is not a criticism of model benchmarks. Aptitude tests are useful — they tell you whether someone is capable of the work. But no HR department would stop at aptitude tests. They would also measure performance on the job. AI evaluation should do the same.
Aptitude test → Certification → Performance review. Three layers. Each measures something different. Each is necessary. None is sufficient alone.
Model benchmark → Output quality → ??? Most enterprises stop at the first two. The third layer — operator performance — is missing.
Operator performance is the performance review.
MO§ES™ is the performance review layer for AI operators. It measures observed behavior in real work — not what operators can do in principle, but what they actually do when interacting with AI.
The method is content-free token telemetry. Every interaction between a human and an AI produces token counts: INPUT, OUTPUT, CACHE READ, CACHE WRITE. These four numbers, tracked over time, reveal:
- How efficiently the operator manages context (Leverage)
- How much productive output they generate relative to total token flow (Yield)
- How much signal they produce relative to noise (Token SNR)
- Whether they build new context or just reuse old context (Construction)
- Whether they are improving over time (Composite Score trend)
No prompts. No outputs. No content inspection. Just the structural patterns of how humans interact with AI — the same way a performance review observes behavior without reading the employee's emails.
What AI evaluation should borrow from human evaluation.
Three principles from a century of human evaluation research that AI evaluation has yet to fully adopt:
What someone can do in a test is not what they will do in the job. Measure both. Do not assume one from the other.
Performance is context-dependent. Evaluate operators in their actual work context, not in a simulated test. MO§ES™ measures real sessions, not test scenarios.
Performance evaluation should route development, not punishment. MO§ES™ labels all results DEVELOPMENTAL — results inform training and intervention, not adverse employment actions.
The operator layer is the performance review AI has been missing.
AI evaluation has built excellent aptitude tests (benchmarks) and output quality checks (evaluation tools). What it has not built — until MO§ES™ — is the performance review: a way to measure how effectively humans actually operate AI in real work.
The lesson from a century of human evaluation research is clear: capability tests are necessary but not sufficient. To understand performance, you have to measure performance. MO§ES™ does that for AI operators, using content-free token telemetry, without inspecting a single prompt.
Common questions.
Why should AI evaluation learn from how we test humans?
Human evaluation has evolved over a century to measure performance in context, not just capability in isolation. Skills tests, certifications, and performance reviews all distinguish between what someone can do and what they actually do. AI evaluation should make the same distinction for operators.
What is the difference between capability and performance in AI evaluation?
Capability is what an AI model or operator can do in principle, measured by tests and benchmarks. Performance is what they actually do in real work, measured by observed behavior. Model benchmarks measure capability. MO§ES™ measures operator performance using content-free token telemetry.
How does MO§ES™ evaluate AI operators?
MO§ES™ measures operators using content-free token telemetry — INPUT, OUTPUT, CACHE READ, and CACHE WRITE counts. These reveal how effectively operators manage context, convert input to output, and improve over time. No prompt or output content is needed.
Go deeper.
The core concept.
Why measuring usage isn't measuring performance.
The structured approaches.