AI agent evaluation, extended to the human operator.
AI agent evaluation assesses how well autonomous and semi-autonomous AI agents perform tasks. But agents do not operate in a vacuum — humans direct them, prompt them, and manage their context. MO§ES™ measures the human side of human-agent collaboration.
What is AI agent evaluation?
AI agent evaluation is the process of assessing how well autonomous and semi-autonomous AI agents perform tasks. It measures goal completion, reasoning quality, tool use, and safety of agent workflows.
Current AI agent evaluation focuses on the agent itself:
- Goal completion — did the agent achieve the objective?
- Reasoning quality — did the agent plan and execute effectively?
- Tool use — did the agent select and use tools correctly?
- Safety — did the agent stay within bounds?
Tools like DeepEval, Braintrust, and Anthropic's eval frameworks address these dimensions. But they evaluate the agent. They do not evaluate the human directing the agent.
Who is evaluating the operator?
AI agents are not autonomous in enterprise settings. They are directed by humans. The quality of that direction — the prompts, the context management, the task decomposition, the feedback loops — determines whether the agent succeeds or fails.
An agent that scores well on benchmarks can fail in production when operated by someone who does not manage context, decompose tasks effectively, or iterate on agent output. The operator is the variable that agent evaluations do not measure.
MO§ES™ measures the operator layer using content-free token telemetry. When a human directs an AI agent, every interaction produces token counts — INPUT, OUTPUT, CACHE READ, CACHE WRITE. These counts reveal:
- How efficiently the operator builds context for the agent
- Whether the operator reuses prior context or starts fresh each time
- How much productive output the agent produces relative to input
- Whether the operator is improving over time
The unit of measurement is the session.
MO§ES™ evaluates human-agent collaboration at the session level. A session is one continuous interaction between an operator and an AI agent. Token telemetry is collected per session and aggregated per operator.
This means you can benchmark operators against each other, track improvement over time, and identify which operators are getting more out of their AI agents — all without reading a single prompt or output.
Go deeper.
The broader concept — all four layers of AI evaluation.
What an AI evaluator is and does — the role and the platform.
The full MO§ES™ evaluation framework.