Early Access Pilot · The Baseline Assessment

One assessment. One baseline. One answer.

The shortest path from skepticism to evidence. We measure your AI operators for 14 days, design one targeted intervention, deploy it, then re-measure for 14 days. If the target metric moves and non-target metrics hold, you have proof. If it doesn't, you have a diagnosis. Either way, you know.

30 DAYS CONTROLLED EVIDENCE-GRADE
The Problem

Everyone measures AI adoption. Nobody measures AI performance.

Your company bought AI licenses. People are using them. Login frequency is up. Token spend is up. The dashboard looks green. But you don't know if anyone is actually getting better at operating AI. And you can't answer the question your board is asking: "Is this investment producing returns, or just activity?"

The Baseline Assessment is designed to answer that question with controlled evidence in 30 days. Not a year-long study. Not a vendor whitepaper. A controlled before-and-after measurement on your own operators, with your own data, producing evidence labeled CAUSATION if the design holds.

The Design

Four phases. Thirty days. One answer.

PHASE 1 · DAYS 1–14
Baseline

Collect canonical token telemetry for 14 days. No intervention. No changes. Establish the baseline for all five canonical metrics (Yield, Leverage, Token SNR, Log Leverage, Construction) across your cohort of 25–100 operators.

Output: Baseline metric profiles, percentile bands, archetype classification, diagnostic flags.

PHASE 2 · DAY 15
Diagnose + Prescribe

On day 15, we analyze the baseline. We identify the weakest metric across the cohort — the one with the most room for improvement. We design one targeted intervention specifically for that weakness.

Output: Diagnostic report, intervention design, target metric selection, non-target metrics to monitor for regression.

PHASE 3 · DAYS 16–28
Intervene

Deploy the intervention. This is not "more AI training." It's a specific, targeted change: a context structuring workshop, an iteration pattern coaching session, a model switching protocol for specific task types, or a workflow redesign. The intervention is declared before measurement resumes.

Output: Intervention deployed, pre-declared measurement plan locked.

PHASE 4 · DAY 29–30
Re-measure + Report

Compare post-intervention metrics against the baseline. The target metric either moved or it didn't. Non-target metrics either held or they didn't. The evidence label is assigned based on the design.

Output: Evidence report with CAUSATION, ASSOCIATION, or INCONCLUSIVE label. Full before/after comparison. Recommended next steps.

What You Get

The deliverable.

Baseline Report

Operator and cohort baselines across all five canonical metrics. Percentile bands, archetype classification, diagnostic flags. Delivered at day 14.

Intervention Design

One targeted intervention designed from your baseline diagnostics. Target metric declared. Non-target metrics monitored. Delivered at day 15.

Evidence Report

Before/after comparison with evidence label: CAUSATION, ASSOCIATION, or INCONCLUSIVE. Full metric deltas. Recommended next steps. Delivered at day 30.

Evidence Grades

What the evidence label means.

CAUSATION

Target metric moved in the expected direction. Non-target metrics held. The intervention caused the change. You can deploy with confidence.

ASSOCIATION

Target metric moved, but non-target metrics also shifted, or the design had a confound. The change is associated with the intervention but not cleanly caused by it.

INCONCLUSIVE

Target metric did not move, or moved ambiguously. You have a diagnosis — the intervention didn't work as designed. That's still valuable evidence.

Parameters

What's included.

  • Cohort size: 25–100 operators
  • Duration: 30 days (14 baseline + 1 intervention + 14 re-measurement)
  • Data: Canonical token telemetry only (no prompt content)
  • Metrics: All 5 canonical (Yield, Leverage, Token SNR, Log Leverage, Construction)
  • Intervention: One targeted intervention, designed from baseline diagnostics
  • Reports: 3 (baseline, intervention design, evidence)
  • Evidence grade: CAUSATION, ASSOCIATION, or INCONCLUSIVE
  • Privacy: Content-free. No prompt text, no output text, no conversation content.
  • Deployment level: Level 1 (Canonical Telemetry) or Level 2 (API Enriched)
  • Price: $15,000
Why This Works

The logic.

Most AI measurement programs fail because they try to measure everything, take a year, and produce a 200-page report that nobody reads. The Baseline Assessment inverts that: measure one thing, change one thing, re-measure one thing, report in 30 days.

The assessment is small enough to execute quickly, controlled enough to produce evidence, and specific enough to answer a real business question: "Can we improve how our people operate AI, and can we prove it?"

If the answer is yes, you have a repeatable playbook. Run it again on a different metric. If the answer is no, you have a diagnosis — and a reason to investigate deeper with an extended engagement.

Governance

What this is not.

  • Not a personnel evaluation. Operators are measured by behavior, not ranked by name.
  • Not a leaderboard. No individual scores are published or shared with management by name.
  • Not surveillance. No prompt content, no output text, no conversation content is collected.
  • Not a productivity claim. We measure operating patterns, not business outcomes. Outcome correlation is labeled ASSOCIATION unless a controlled experiment supports CAUSATION.
  • Not a one-size-fits-all benchmark. The intervention is designed from your baseline, not a generic playbook.
Start

Run the Baseline Assessment.

If you've deployed AI and can't answer whether your operators are getting better, this is the fastest way to find out. 30 days. One assessment. One answer.

Book a Demo Run the Demo
Related

Go deeper