How to Engage

From first contact to measured outcomes. The complete pilot process.

The Process

Five steps from contact to readout.

Discovery Call Baseline Diagnose Intervene Re-evaluate Readout

Step 1 — Discovery Call

Email pilots@mos2es.org to schedule a 30-minute call. We discuss:

  • Team size and structure (25–100 users for baseline)
  • AI tools currently in use (Claude, Cursor, Copilot, Windsurf, etc.)
  • What you want to measure (operator performance, workflow efficiency, intervention impact)
  • Privacy and governance requirements

No commitment required. The call is to determine fit.

Step 2 — 30-Day Baseline Assessment

We collect content-free token telemetry from your team's AI tool usage. This means:

  • What we collect: input tokens, output tokens, cache read, cache write
  • What we do NOT collect: prompt text, response content, code content, or any personal data
  • How long: 30 days of passive observation
  • What you get: baseline metrics across 8 canonical measures

The 8 canonical metrics computed from your baseline:

Yield (Υ)

(cache_read × output) / input² — the signature metric. How much usable output you get relative to what you put in.

Leverage

cache_read / input — how efficiently you reuse prior context.

Token SNR

output / (input + output) — signal-to-noise ratio. What fraction of total tokens are useful output.

Log Leverage (10xDEV)

log₁₀(cache_read / input) — leverage on a logarithmic scale. 10x = one order of magnitude.

Construction

cache_write / cache_read — how much new context you build vs reuse.

Velocity

output / input — raw output efficiency. Tokens out per token in.

Scale V

log₁₀(input + output + cache_write + cache_read) — total volume on a log scale.

Efficiency

(cache_read + cache_write + output) / input / 4 — composite display metric.

Composite scores are DEVELOPMENTAL, never PERSONNEL. No punitive use, no employee leaderboard.

Step 3 — Bespoke Evaluation Design

We design evaluations around your specific workflows, not generic benchmarks. This includes:

  • Workflow-specific operator evaluation (your tools, your tasks)
  • Internal benchmarking (how do your operators compare to each other?)
  • External reference comparisons (how do your operators compare to the public SigRank field?)
  • Capability gap diagnosis (where are the bottlenecks?)

Step 4 — Intervention Testing

Based on the diagnosis, we identify targeted interventions and test them:

  • Training changes (prompt patterns, context management strategies)
  • Tool configuration changes (model selection, context window settings)
  • Workflow restructuring (task assignment, collaboration patterns)

We re-measure after the intervention period to determine what worked. All outcomes are labeled ASSOCIATION, never CAUSATION, unless backed by controlled experiments.

Step 5 — Pilot Readout

You receive a full pilot readout that includes:

  • Baseline metrics and percentile bands
  • Internal and external benchmark comparisons
  • Diagnostic findings and capability gap analysis
  • Intervention results (what worked, what didn't)
  • Recommended next steps for ongoing measurement
Escalation

Human escalation paths.

Pilot Inquiries

pilots@mos2es.org
For enterprise pilot discussions, scheduling, and process questions.

MCP Write Operations

The MCP server has 5 write tools that require authorization. Contact us to provision API credentials for automated pilot data submission.

Technical Questions

Read the docs or connect to the MCP server for technical integration questions.

General Contact

Contact form for any other inquiries.

Ready to start?

Email pilots@mos2es.org or review the 30-day baseline assessment details.