AI operator benchmarking.
A performative benchmark compares how operators actually perform under observed or defined operating conditions — rather than relying on self-reported proficiency, certification, or generic knowledge tests. MO§ES™ defines 13 benchmark classes and a selection algorithm that chooses the right benchmark for each evaluation question.
What a performative benchmark is.
A performative benchmark measures observed operator behavior against defined conditions. It does not ask operators what they know — it measures what they do. The benchmark is constructed from real telemetry, real workflows, and real operating conditions.
Every benchmark is defined by four axes: the task being performed, the model being used, the operator being measured, and the context in which the work happens. Changing any axis changes the benchmark. A high-performing operator in one context may not be the strongest in another. "Best operator" should require the same qualifier as "best model." Best at what?
Benchmark types and their uses.
MO§ES™ defines 13 benchmark classes organized into three tiers: standard metrics, contextual comparisons, and temporal analyses. Each class answers a different evaluation question.
Common canonical metrics usable across organizations without customization.
Compare operators and teams within the organization.
Compare against relevant reference populations (SigRank field).
Built around organization-specific workflows, roles, and tasks.
Compare performance over time across measurement windows.
Compare pre- and post-intervention performance.
Benchmarks scoped to a specific role or job function.
Benchmarks scoped to a specific AI model or provider.
Benchmarks scoped to a specific task type or workflow.
Compare operator performance across different AI providers.
Aggregate operator metrics to the team level for organizational comparison.
Position operators in percentile bands within the reference cohort.
Benchmark the gap between usage rank and evaluation rank.
How the right benchmark is chosen.
Not every evaluation question requires every benchmark class. The selection algorithm matches the evaluation question to the appropriate benchmark classes based on the comparison scope, temporal requirements, and available data.
The algorithm evaluates four factors:
- Comparison scope — internal only, external field, or both
- Temporal requirement — single window or longitudinal
- Customization level — standard metrics or bespoke workflow
- Intervention context — baseline only or pre/post comparison
For example, a question about whether an operator is improving over time selects the Longitudinal and Percentile Band classes. A question about whether a team is outperforming the organizational average selects the Internal Cohort and Team-Level classes. A question about whether a specific workflow change improved efficiency selects the Intervention and Bespoke Workflow classes.
How results are reported.
Benchmark results are reported as percentile bands, not absolute scores. This normalizes across cohorts of different sizes and skill distributions.
Benchmarks should be contextual. A high-performing operator in one workflow may not be the strongest in another. The system never collapses benchmark, diagnosis, and validated outcome into a single claim.
What benchmarking does not tell you.
Benchmark results are labeled DEVELOPMENTAL. They provide structural context, not a universal ranking of "better" or "worse" operators.
- Not a universal ranking. Benchmarks compare operators under specific conditions. Results are contextual, not absolute.
- Not a performance verdict. A high percentile rank means the operator scores well on the measured metrics under the defined conditions. It does not mean they are a better employee or more productive worker.
- Not an employee ranking. DEVELOPMENTAL labels mean results route workflows, not people. No adverse employment actions are permitted from pilot data.
- Reference population is structural. External field comparison provides structural context, not a universal standard of operator quality.
- Validation required. The relationship between benchmark position and validated business outcomes has not been independently established. Correlations are labeled ASSOCIATION, never CAUSATION.
Every benchmark result carries provenance: benchmark class, reference cohort, percentile band, evidence label, decision-use label (DEVELOPMENTAL), and synthetic-data flag where applicable.
Read alongside.
The 0–100 index that benchmark positions are based on.
Benchmark class 13 — the gap between usage and evaluation rank.
Benchmark class 6 — pre/post intervention comparison.