Concept · Benchmarking

AI operator benchmarking.

A performative benchmark compares how operators actually perform under observed or defined operating conditions — rather than relying on self-reported proficiency, certification, or generic knowledge tests. MO§ES™ defines 13 benchmark classes and a selection algorithm that chooses the right benchmark for each evaluation question.

MEASURED DERIVED DEVELOPMENTAL
Definition

What a performative benchmark is.

A performative benchmark measures observed operator behavior against defined conditions. It does not ask operators what they know — it measures what they do. The benchmark is constructed from real telemetry, real workflows, and real operating conditions.

TASK × MODEL × OPERATOR × CONTEXT = PERFORMATIVE BENCHMARK

Every benchmark is defined by four axes: the task being performed, the model being used, the operator being measured, and the context in which the work happens. Changing any axis changes the benchmark. A high-performing operator in one context may not be the strongest in another. "Best operator" should require the same qualifier as "best model." Best at what?

The 13 Benchmark Classes

Benchmark types and their uses.

MO§ES™ defines 13 benchmark classes organized into three tiers: standard metrics, contextual comparisons, and temporal analyses. Each class answers a different evaluation question.

1. Standard Metrics

Common canonical metrics usable across organizations without customization.

2. Internal Cohort

Compare operators and teams within the organization.

3. External Field

Compare against relevant reference populations (SigRank field).

4. Bespoke Workflow

Built around organization-specific workflows, roles, and tasks.

5. Longitudinal

Compare performance over time across measurement windows.

6. Intervention

Compare pre- and post-intervention performance.

7. Role-Specific

Benchmarks scoped to a specific role or job function.

8. Model-Specific

Benchmarks scoped to a specific AI model or provider.

9. Task-Specific

Benchmarks scoped to a specific task type or workflow.

10. Cross-Model

Compare operator performance across different AI providers.

11. Team-Level

Aggregate operator metrics to the team level for organizational comparison.

12. Percentile Band

Position operators in percentile bands within the reference cohort.

13. Divergence

Benchmark the gap between usage rank and evaluation rank.

Selection Algorithm

How the right benchmark is chosen.

Not every evaluation question requires every benchmark class. The selection algorithm matches the evaluation question to the appropriate benchmark classes based on the comparison scope, temporal requirements, and available data.

DEFINE QUESTION IDENTIFY SCOPE SELECT CLASSES COMPUTE BENCHMARK REPORT WITH BANDS

The algorithm evaluates four factors:

  • Comparison scope — internal only, external field, or both
  • Temporal requirement — single window or longitudinal
  • Customization level — standard metrics or bespoke workflow
  • Intervention context — baseline only or pre/post comparison

For example, a question about whether an operator is improving over time selects the Longitudinal and Percentile Band classes. A question about whether a team is outperforming the organizational average selects the Internal Cohort and Team-Level classes. A question about whether a specific workflow change improved efficiency selects the Intervention and Bespoke Workflow classes.

Percentile Bands

How results are reported.

Benchmark results are reported as percentile bands, not absolute scores. This normalizes across cohorts of different sizes and skill distributions.

Median
50th
Top 25%
75th
Top 10%
90th
Top 5%
95th
Top 1%
99th
Top 0.1%
99.9th

Benchmarks should be contextual. A high-performing operator in one workflow may not be the strongest in another. The system never collapses benchmark, diagnosis, and validated outcome into a single claim.

Governance Caveats

What benchmarking does not tell you.

Benchmark results are labeled DEVELOPMENTAL. They provide structural context, not a universal ranking of "better" or "worse" operators.

  • Not a universal ranking. Benchmarks compare operators under specific conditions. Results are contextual, not absolute.
  • Not a performance verdict. A high percentile rank means the operator scores well on the measured metrics under the defined conditions. It does not mean they are a better employee or more productive worker.
  • Not an employee ranking. DEVELOPMENTAL labels mean results route workflows, not people. No adverse employment actions are permitted from pilot data.
  • Reference population is structural. External field comparison provides structural context, not a universal standard of operator quality.
  • Validation required. The relationship between benchmark position and validated business outcomes has not been independently established. Correlations are labeled ASSOCIATION, never CAUSATION.

Every benchmark result carries provenance: benchmark class, reference cohort, percentile band, evidence label, decision-use label (DEVELOPMENTAL), and synthetic-data flag where applicable.

Related Concepts

Read alongside.

The 0–100 index that benchmark positions are based on.

Benchmark class 13 — the gap between usage and evaluation rank.

Benchmark class 6 — pre/post intervention comparison.

Read the Methodology Request a Pilot