Concept · Composite Metric

AI Operator Development Index.

The AI Operator Development Index is a 0–100 composite score that combines the four primary canonical metrics into a single summary of operator behavior. It is labeled DEVELOPMENTAL — a structural signal for routing workflows, not a performance verdict or an employee ranking.

DERIVED DEVELOPMENTAL VALIDATION REQUIRED
Definition

The composition.

The index normalizes each of the four primary canonical metrics to a 0–100 scale within the operator's cohort, then combines them with fixed weights. The weights reflect the relative importance of each metric in the current framework design.

ODI = (Leverage × 0.30) + (Yield × 0.30) + (Token SNR × 0.20) + (Construction × 0.20)
ComponentWeightWhat it captures
Leverage30%Context reuse and construction efficiency relative to new input
Yield30%Productive output share of total token flow
Token SNR20%Signal quality — output relative to input and reused context
Construction20%Ratio of new context built to context reused

Each component is percentile-normalized within the cohort before weighting. This means an operator at the 90th percentile on Leverage contributes 90 × 0.30 = 27 points from that component. The maximum possible score is 100, representing an operator at the 99th+ percentile on all four metrics. The minimum is 0, representing an operator at the bottom of the cohort on all four.

Interpretation

What the index tells you.

The index is a summary, not a substitute for the individual metrics. It provides a single number for quick comparison and ranking, but it obscures the component-level patterns that drive the score.

High Index

The operator scores well across multiple canonical metrics. This may indicate efficient context management, high output efficiency, good signal quality, and balanced context construction. It does not prove better outcomes — it describes a structural operating pattern.

Low Index

The operator scores below the cohort median across multiple metrics. This may indicate input-heavy patterns, low output efficiency, or context management styles that do not align with the current weighting. It is not a performance verdict.

Two operators with the same index score can have very different component profiles. One may score high on Leverage and Yield but low on SNR and Construction. Another may score evenly across all four. The index does not distinguish between these profiles — that requires reading the component metrics individually.

How MO§ES™ Uses It

The index in the evaluation pipeline.

The index is computed after all four component metrics have been measured and benchmarked. It enters the pipeline at two stages.

Ranking

Operators are ranked by the index within their cohort. Percentile bands show where each operator falls in the composite distribution.

Tracking

The index is tracked over time to show movement. An operator's index can rise, fall, or remain stable across measurement windows.

Divergence

The index is compared against usage volume to identify operators whose composite score does not match their usage rank — the divergence signal.

In the synthetic demo cohort of 50 operators across 1,668 observations spanning a 30-day window, the index distribution is broad with a clear right tail. The top 10% of operators score above 75 on the index, while the bottom 10% score below 35. These are structural observations from synthetic data, not validated performance findings.

Governance Caveats

What the index does not tell you.

The AI Operator Development Index is labeled DEVELOPMENTAL. It is a structural summary whose relationship to actual performance outcomes is still being tested.

  • Not a performance score. The index summarizes structural operating behavior. It does not measure output quality, correctness, business value, or productivity.
  • Not an employee ranking. DEVELOPMENTAL labels mean results route workflows, not people. No adverse employment actions are permitted from pilot data. The index is never used for automatic adverse actions.
  • Weights are design choices, not validated truths. The 30/30/20/20 weighting reflects the current framework design. Alternative weightings would produce different rankings. The weights have not been independently validated against outcome measures.
  • Obscures component detail. Two operators with the same index can have very different operating profiles. The index is a summary, not a diagnosis.
  • Validation required. The relationship between the index and validated business outcomes has not been independently established. Correlations are labeled ASSOCIATION, never CAUSATION.

Every index value carries provenance: source telemetry window, operator cohort, component scores, evidence label (DERIVED), decision-use label (DEVELOPMENTAL), and synthetic-data flag where applicable.

Related Concepts

Read alongside.

30% of the index. (R + W) / I.

30% of the index. O / (I + O + R + W).

20% of the index. O / (I + O + R).

20% of the index. W / R.

Where index rank does not match usage rank.

DEVELOPMENTAL, HYPOTHESIS, ASSOCIATION labels.

Read the Methodology Request a Pilot