Commercial Pilot · The Killer Experiment

One experiment. One answer.

The shortest path from skepticism to evidence. We measure your AI operators for 14 days, design one targeted intervention, deploy it, then re-measure for 14 days. If the target metric moves and non-target metrics hold, you have proof. If it doesn't, you have a diagnosis. Either way, you know.

30 DAYS CONTROLLED EVIDENCE-GRADE
The Problem

Everyone measures AI adoption. Nobody measures AI performance.

Your company bought AI licenses. People are using them. Login frequency is up. Token spend is up. The dashboard looks green. But you don't know if anyone is actually getting better at operating AI. And you can't answer the question your board is asking: "Is this investment producing returns, or just activity?"

The killer experiment is designed to answer that question with controlled evidence in 30 days. Not a year-long study. Not a vendor whitepaper. A controlled before-and-after measurement on your own operators, with your own data, producing evidence labeled CAUSATION if the design holds.

The Design

Four phases. Thirty days. One answer.

PHASE 1 · DAYS 1–14
Baseline

Collect canonical token telemetry for 14 days. No intervention. No changes. Establish the baseline for all five canonical metrics (Yield, Leverage, Token SNR, Log Leverage, Construction) across your cohort of 25–100 operators.

Output: Baseline metric profiles, percentile bands, archetype classification, diagnostic flags.

PHASE 2 · DAY 15
Diagnose + Prescribe

On day 15, we analyze the baseline. We identify the weakest metric across the cohort — the one with the most room for improvement. We design one targeted intervention specifically for that weakness.

Output: Diagnostic report, intervention design, target metric selection, non-target metrics to monitor for regression.

PHASE 3 · DAYS 16–28
Intervene

Deploy the intervention. This is not "more AI training." It's a specific, targeted change: a context structuring workshop, an iteration pattern coaching session, a model switching protocol for specific task types, or a workflow redesign. The intervention is declared before measurement resumes.

Output: Intervention deployed, pre-declared measurement plan locked.

PHASE 4 · DAYS 16–30
Re-measure

Continue collecting telemetry for 14 days post-intervention. Compare against the baseline. Look at the target metric (did it improve?) and non-target metrics (did anything regress?). If the target moved and nothing else broke, you have controlled evidence.

Output: Final report with evidence grade, metric deltas, regression check, and a clear verdict: the intervention worked, didn't work, or produced trade-offs.

What You Get

The deliverable.

REPORT 1
Baseline Report

Day 15. Full cohort baseline: metric profiles, percentile bands, archetype distribution, diagnostic flags, weakest-metric identification, intervention recommendation.

REPORT 2
Intervention Design

Day 15. The specific intervention, target metric, expected effect direction, non-target metrics to monitor, and pre-declared measurement plan. Locked before deployment.

REPORT 3
Evidence Report

Day 30. Metric deltas (target + non-target), evidence grade (ASSOCIATION or CAUSATION), regression check, statistical significance, and a clear verdict.

Evidence Grades

What the evidence label means.

GradeMeaningWhen Applied
CAUSATIONThe intervention caused the metric change. Controlled before-and-after design with no regression in non-target metrics.Target metric moved significantly, non-target metrics held.
ASSOCIATIONThe metric change is associated with the intervention but confounding factors cannot be ruled out.Target metric moved but non-target metrics also changed, or sample size is small.
INCONCLUSIVENo significant metric change detected.Target metric did not move, or moved within noise threshold.

We label evidence honestly. If the design doesn't support a causal claim, we say ASSOCIATION. If nothing moved, we say INCONCLUSIVE. We don't dress up null results.

Parameters

What's included.

  • Cohort size: 25–100 operators
  • Duration: 30 days (14 baseline + 1 intervention + 14 re-measurement)
  • Data: Canonical token telemetry only (no prompt content)
  • Metrics: All 5 canonical (Yield, Leverage, Token SNR, Log Leverage, Construction)
  • Intervention: One targeted intervention, designed from baseline diagnostics
  • Reports: 3 (baseline, intervention design, evidence)
  • Evidence grade: CAUSATION, ASSOCIATION, or INCONCLUSIVE
  • Privacy: Content-free. No prompt text, no output text, no conversation content.
  • Deployment level: Level 1 (Canonical Telemetry) or Level 2 (API Enriched)
  • Price: $15,000
Why This Works

The logic.

Most AI measurement programs fail because they try to measure everything, take a year, and produce a 200-page report that nobody reads. The killer experiment inverts that: measure one thing, change one thing, re-measure one thing, report in 30 days.

The experiment is small enough to execute quickly, controlled enough to produce evidence, and specific enough to answer a real business question: "Can we improve how our people operate AI, and can we prove it?"

If the answer is yes, you have a repeatable playbook. Run it again on a different metric. If the answer is no, you have a diagnosis — and a reason to investigate deeper with a full pilot.

Governance

What this is not.

  • Not a personnel evaluation. Operators are measured by behavior, not ranked by name.
  • Not a leaderboard. No individual scores are published or shared with management by name.
  • Not surveillance. No prompt content, no output text, no conversation content is collected.
  • Not a productivity claim. We measure operating patterns, not business outcomes. Outcome correlation is labeled ASSOCIATION unless a controlled experiment supports CAUSATION.
  • Not a one-size-fits-all benchmark. The intervention is designed from your baseline, not a generic playbook.
Start

Run the killer experiment.

If you've deployed AI and can't answer whether your operators are getting better, this is the fastest way to find out. 30 days. One experiment. One answer.

Book the Killer Experiment See Full Pilot Options
Related

Go deeper