Three ways to engage.
From a 30-day proof to a year-round operating index. Each proposition uses the same content-free measurement engine, the same five canonical metrics, and the same governance rules. The difference is depth, duration, and how much of your AI operating surface you want to understand.
The Killer Experiment — $15,000
For the skeptical buyer who needs proof before commitment. One controlled experiment. 30 days. One cohort, one intervention, one re-measurement. If the target metric moves and non-target metrics hold, you have controlled evidence. If it doesn't, you have a diagnosis.
What's included:
- 14-day baseline measurement (all 5 canonical metrics)
- Diagnostic report + intervention design (day 15)
- One targeted intervention deployment
- 14-day re-measurement
- Evidence report with CAUSATION, ASSOCIATION, or INCONCLUSIVE grade
- Content-free telemetry (no prompt text)
Best for:
Head of AI, CIO, or Transformation leader who has deployed AI but can't answer whether operators are improving. The fastest path from "we think it's working" to "we have evidence."
Learn More Book the ExperimentThe Enterprise Pilot — $45,000
For the committed buyer who wants the full diagnostic surface. 90 days. Up to 500 operators. All 15 eval families. Full diagnostic, intervention, and re-measurement loop. Choose from 12 pilot products tailored to your business question.
What's included:
- 30-day baseline measurement (all 5 canonical metrics + cohort distributions)
- Full diagnostic suite: usage vs operation divergence, context architecture, longitudinal movement, platform/model sensitivity, cohort composition, workflow fit, team composition, capability dependency risk, org AI topology, operator similarity
- Up to 3 targeted interventions with pre-declared measurement plans
- 60-day re-measurement window
- Executive brief + full technical report
- Production gates (GATE-001/002/003) for developmental routing
- Optional: outcome correlation join (ASSOCIATION grade)
- Content-free telemetry (no prompt text)
Choose your pilot product:
| Pilot | Question | Best Buyer |
|---|---|---|
| Pilot 1 AI Workforce Baseline | What does our AI workforce look like? | Head of AI |
| Pilot 2 Capability Distribution | Do we have org capability or a few power users? | Head of AI / BU |
| Pilot 4 Training Evaluation | Did our training change how people operate AI? | L&D |
| Pilot 5 Model/Tool Evaluation | What changes when we switch models? | CIO |
| Pilot 7 Workflow Diagnostic | Is the constraint the operator, tool, or workflow? | Operations |
| Pilot 8 Team Comparison | Why do same-stack teams operate differently? | Transformation |
| Pilot 11 Meta-Pilot | Does this metric predict our KPIs? | CIO |
Best for:
Head of AI, CIO, Transformation, or L&D leader who needs a comprehensive diagnostic of how their organization operates AI. Not just "is it working?" but "where is it working, where isn't, and why?"
See Full Pilot Details Book the PilotThe Annual Operating Index — $150,000/year
For the enterprise that treats AI operator performance as an ongoing management discipline, not a one-time study. Year-round measurement. Monthly reporting. Ongoing intervention testing. Governed experiment framework. The full MO§ES™ engine as a managed service.
What's included:
- Continuous canonical telemetry collection (all operators, all sessions)
- Monthly operating index report: cohort distributions, metric trends, band movement, learning curves, dependency risk, topology changes
- Quarterly diagnostic deep-dive: full 15-eval-family analysis
- Up to 6 governed experiments per year (EVAL-012 Experiment as Product)
- Unlimited interventions with pre-declared measurement plans
- Production gates with custom thresholds and routing rules
- Outcome correlation joins (ASSOCIATION grade, with optional controlled experiments for CAUSATION)
- Executive dashboard with monthly delivery review
- Dedicated measurement engineer + quarterly business review
- API access to all metrics and reports
- MCP server integration for AI agent access to operating data
- Content-free telemetry (no prompt text), Level 3 governed deployment
Best for:
CIO, Head of AI, or Transformation leader who has moved beyond "is AI working?" to "how do we continuously improve how our organization operates AI?" The operating index becomes a management discipline, not a project.
Book a Demo See the Full ProductWhich proposition fits?
| Killer Experiment | Enterprise Pilot | Annual Index | |
|---|---|---|---|
| Price | $15,000 | $45,000 | $150,000/yr |
| Duration | 30 days | 90 days | 12 months |
| Operators | 25–100 | Up to 500 | Unlimited |
| Eval families | 1 (EVAL-007) | All 15 | All 15 + continuous |
| Interventions | 1 | Up to 3 | Unlimited |
| Reports | 3 | Executive + technical | Monthly + quarterly |
| Experiments | 1 controlled | Up to 3 | 6/year governed |
| Deployment level | Level 1 | Level 1–2 | Level 3 governed |
| Evidence grade | CAUSATION possible | ASSOCIATION + CAUSATION | ASSOCIATION + CAUSATION |
| Best buyer | Skeptical CIO | Head of AI | CIO / Transformation |
| Question answered | "Can we prove it works?" | "Where is it working?" | "How do we keep improving?" |
What's true across all three.
- Content-free. No prompt text, no output text, no conversation content. Four token counts plus structural signals.
- Five canonical metrics. Yield, Leverage, Token SNR, Log Leverage, Construction. Derived from the same four token numbers.
- No leaderboards. Operators are measured by behavior, not ranked by name. No punitive labels.
- Evidence labeled honestly. ASSOCIATION by default. CAUSATION only when a controlled experiment supports it.
- Developmental, not personnel. Composite scores route coaching and intervention, not hiring decisions.
- Privacy by design. Aligns with GDPR, SOC 2, and enterprise data governance.
Pick your proposition.
Not sure which one fits? Start with the Killer Experiment. If it produces evidence, you'll know the engine works. Then upgrade to the Pilot or Annual Index.
Talk to Us Try the Demo First