MO§ES™ vs ccusage
ccusage is a CLI tool that reports token usage for Claude Code sessions — how many tokens you consumed, what it cost, which models you used. MO§ES™ measures how well you used those tokens — operator performance, leverage, yield, and benchmarking across all AI systems. The distinction is usage reporting vs operator evaluation.
Explore the 30-Day Pilot See the MethodologyDifferent objects, different methods.
ccusage reports usage. MO§ES™ evaluates performance. ccusage's question is "how many tokens did you use in Claude Code, and what did it cost?" MO§ES™'s question is "how effectively did you operate AI systems to accomplish your work?" Both consume token data. They produce fundamentally different outputs.
ccusage parses Claude Code's local session logs, extracts token counts and cost estimates, and presents them as reports — daily totals, session breakdowns, model-level summaries. The object is the usage. The method is log parsing. The output is a usage report.
MO§ES™ captures the canonical telemetry surface — INPUT, OUTPUT, CACHE READ, CACHE WRITE — across all AI systems the operator uses, derives five performance metrics from those signals, and benchmarks the operator against a population. The object is the operator's performance. The method is telemetry-based measurement. The output is a performance profile.
Comparison at a glance.
| Dimension | MO§ES™ | ccusage |
|---|---|---|
| What it measures | AI operator performance across all AI systems via content-free token telemetry | Claude Code token usage and cost estimates per session and per day |
| Mechanism | Canonical telemetry (INPUT, OUTPUT, CACHE READ, CACHE WRITE) → 5 derived metrics → benchmarks | CLI tool that parses Claude Code session logs for token counts |
| Scope | All AI systems, all providers, all roles, all workflows | Claude Code only; single tool, single provider |
| Governance | DEVELOPMENTAL gates, ASSOCIATION never CAUSATION, HYPOTHESIS never fact, provenance on every measurement | Usage reporting; open-source CLI utility |
| Pricing | 30-day enterprise pilot; 6 commercial packages | Free and open-source |
| Best for | Enterprises that need to measure and improve how operators actually perform with AI across all tools | Claude Code users who want to see their token usage and cost estimates |
How many tokens is not how well you used them.
ccusage tells you how many tokens you consumed. This is useful information — it helps you understand your Claude Code usage patterns, estimate costs, and track consumption over time. But token count is a volume metric. It tells you how much you used. It does not tell you how well you used it.
MO§ES™ tells you how effectively you used tokens. Leverage measures how much output you generated per unit of input. Yield measures how much of that output survived into your final work. Token SNR measures the signal-to-noise ratio in your token stream. These are performance metrics — they evaluate quality, not quantity. An operator who used 50,000 tokens with high Leverage and high Yield outperformed an operator who used 200,000 tokens with low Leverage and low Yield. ccusage sees the first operator as "light usage" and the second as "heavy usage." MO§ES™ sees the first as effective and the second as wasteful.
The difference matters because volume-based metrics can be actively misleading. A team that increases its token usage might be getting more productive — or it might be getting less efficient. ccusage cannot distinguish between the two. It reports the increase as a number. MO§ES™ reports whether the increase corresponds to improved or degraded performance. The same data, interpreted through different frameworks, produces opposite conclusions.
One tool vs the entire operating surface.
ccusage is scoped to Claude Code. It parses Claude Code's session logs and reports on Claude Code's token usage. This is a deliberate design choice — ccusage is a focused utility that does one thing well. But it means ccusage sees only a slice of the operator's AI activity.
MO§ES™ is scoped to all AI systems the operator uses. In the demo dataset, measurement spanned 5 AI providers, 13 benchmark classes, and 27 MCP tools. An operator who uses Claude Code for coding, ChatGPT for drafting, and Gemini for research appears in ccusage as only their Claude Code activity. In MO§ES™, they appear as a complete operator with a performance profile that spans every AI system they touch. The operating surface is the full picture, not one tool within it.
This scope difference has practical consequences for enterprise evaluation. If an enterprise wants to know how well its operators use AI, it needs to see across all the AI tools those operators use — not just one. A team that performs well in Claude Code but poorly in other AI systems has a performance gap that ccusage cannot detect. MO§ES™ detects it because it measures the operator, not the tool.
A usage report is not a performance evaluation.
ccusage produces reports — summaries of what happened. "You used 1.2M tokens today, costing $4.80, across 8 sessions." Reports are descriptive. They tell you what occurred. They do not tell you whether what occurred was good, bad, above average, or below average. There is no benchmark, no comparison, no evaluation.
MO§ES™ produces evaluations — performance profiles positioned within a benchmark population. "Your Leverage is 2.8, above the role benchmark of 2.1. Your Yield is 0.38, below the role benchmark of 0.45." Evaluations are comparative and judgmental. They tell you not just what happened but how it ranks, where the gaps are, and what the target is. In the demo dataset, 50 operators were benchmarked across 13 classes, producing a performance field that positioned each operator relative to their peers.
The difference between a report and an evaluation is the difference between data and insight. A usage report gives you a number. An evaluation gives you a number plus a context plus a direction. ccusage gives you the number. MO§ES™ gives you the number, the context, and the direction — with the governance discipline to label every conclusion as DEVELOPMENTAL, every association as never causation, and every claim as hypothesis until proven otherwise.
When to choose which.
- You use Claude Code and want to see your token usage
- You want a free, lightweight CLI tool for usage stats
- You need cost estimates for Claude Code sessions
- Your question is "how many tokens did I use in Claude Code?"
- You want personal usage reporting, not enterprise evaluation
- You need to measure how operators perform with AI across all tools
- You want performance metrics (Leverage, Yield, Token SNR), not usage totals
- You need to benchmark operators against a population and role benchmarks
- You want governance-guardrailed measurement (DEVELOPMENTAL, ASSOCIATION, HYPOTHESIS)
- Your question is "how well are our people operating AI across all systems?"