AI evaluation frameworks, from models to operators.
AI evaluation frameworks provide structured approaches to assessing AI performance. Most frameworks evaluate models or outputs. MO§ES™ is an AI evaluation framework for the operator layer — the human side of AI performance.
What are AI evaluation frameworks?
AI evaluation frameworks are structured methodologies for assessing AI system performance. They define what to measure, how to measure it, how to interpret results, and how to make decisions based on findings.
The major AI evaluation frameworks in use today:
- NIST AI RMF — Risk management framework for AI systems. Focuses on safety, fairness, transparency, and accountability.
- OpenAI Evals — Framework for evaluating LLM capabilities against custom test sets.
- DeepEval — Open-source framework for evaluating LLM outputs using metrics like faithfulness, answer relevance, and context recall.
- Braintrust — Evaluation framework for AI applications with custom evaluators and human review.
- Anthropic evals — Framework for evaluating AI agents on safety, capability, and alignment.
- MO§ES™ — Evaluation framework for the operator layer. Measures how humans use AI using content-free token telemetry.
An evaluation framework for operators.
MO§ES™ is a structured evaluation framework that defines what to measure at the operator layer, how to measure it, how to interpret it, and how to act on it.
Five canonical derived metrics: Yield, Leverage, Token SNR, Construction, and Composite Score. All computed from four token primitives.
Content-free token telemetry. INPUT, OUTPUT, CACHE READ, CACHE WRITE counts per session. No prompt or output content needed.
Cohort benchmarking with percentile bands. Operators compared against peers in the same workflow and model conditions. Divergence scores flag outliers.
Diagnostic patterns route to targeted interventions. DEVELOPMENTAL labels ensure results route workflows, not personnel actions.
Every measurement carries provenance: source telemetry, cohort, evidence label (DERIVED), decision-use label (DEVELOPMENTAL), and synthetic-data flag.
Relationship to business outcomes labeled ASSOCIATION, not CAUSATION, unless validated through controlled experiments.
Go deeper.
The full MO§ES™ evaluation framework.
The core concept.
The governance framework behind MO§ES™.