Concept · Framework

AI evaluation frameworks, from models to operators.

AI evaluation frameworks provide structured approaches to assessing AI performance. Most frameworks evaluate models or outputs. MO§ES™ is an AI evaluation framework for the operator layer — the human side of AI performance.

Definition

What are AI evaluation frameworks?

AI evaluation frameworks are structured methodologies for assessing AI system performance. They define what to measure, how to measure it, how to interpret results, and how to make decisions based on findings.

The major AI evaluation frameworks in use today:

  • NIST AI RMF — Risk management framework for AI systems. Focuses on safety, fairness, transparency, and accountability.
  • OpenAI Evals — Framework for evaluating LLM capabilities against custom test sets.
  • DeepEval — Open-source framework for evaluating LLM outputs using metrics like faithfulness, answer relevance, and context recall.
  • Braintrust — Evaluation framework for AI applications with custom evaluators and human review.
  • Anthropic evals — Framework for evaluating AI agents on safety, capability, and alignment.
  • MO§ES™ — Evaluation framework for the operator layer. Measures how humans use AI using content-free token telemetry.
The MO§ES™ Framework

An evaluation framework for operators.

MO§ES™ is a structured evaluation framework that defines what to measure at the operator layer, how to measure it, how to interpret it, and how to act on it.

What to measure

Five canonical derived metrics: Yield, Leverage, Token SNR, Construction, and Composite Score. All computed from four token primitives.

How to measure

Content-free token telemetry. INPUT, OUTPUT, CACHE READ, CACHE WRITE counts per session. No prompt or output content needed.

How to interpret

Cohort benchmarking with percentile bands. Operators compared against peers in the same workflow and model conditions. Divergence scores flag outliers.

How to act

Diagnostic patterns route to targeted interventions. DEVELOPMENTAL labels ensure results route workflows, not personnel actions.

Governance

Every measurement carries provenance: source telemetry, cohort, evidence label (DERIVED), decision-use label (DEVELOPMENTAL), and synthetic-data flag.

Outcomes

Relationship to business outcomes labeled ASSOCIATION, not CAUSATION, unless validated through controlled experiments.

Related

Go deeper.

The full MO§ES™ evaluation framework.

The core concept.

The governance framework behind MO§ES™.

See Competitor Comparisons Request a Pilot