Competitor Comparison

MO§ES™ vs Vals AI

Vals AI is a business-model analogue: a public benchmark authority and private evaluation infrastructure for AI models and applications. MO§ES™ evaluates human operators, not models. Vals AI is a reference model, not a direct competitor — but the parallel is instructive.

Explore the 30-Day Pilot See the Methodology
The Core Distinction

Evaluate the model or evaluate the operator?

Native analytics measure usage. Skills systems measure predefined capability. Engineering systems measure engineering work. MO§ES™ measures the operator operating technology and builds upward from that object.

Vals AI evaluates AI models and applications. It provides public benchmark authority — independent, credible rankings of how models perform — and private evaluation infrastructure that enterprises can use to test models and apps in their own contexts. This is model evaluation: how good is the AI system?

MO§ES™ evaluates human operators. It measures how people perform when they use AI systems. This is operator evaluation: how good is the person using the AI?

These are complementary, not competitive. An enterprise needs both. It needs to know which models are best (Vals AI) and it needs to know which operators use those models best (MO§ES™). A great model in the hands of a low-leverage operator produces less value than a good model in the hands of a high-leverage operator. The model and the operator are both part of the system, and both need evaluation.

Side by Side

Comparison at a glance.

DimensionMO§ES™Vals AI
What it measuresHuman operator performance via content-free token telemetry across real tasksAI model and application performance via benchmarks and evaluation infrastructure
MechanismCanonical telemetry (INPUT, OUTPUT, CACHE READ, CACHE WRITE) → derived metrics → benchmarksPublic benchmark authority + private evaluation infrastructure for models and apps
ScopeAll human operators across all roles, all AI systems, all workflowsAI models and applications across tasks and domains
GovernanceDEVELOPMENTAL gates, ASSOCIATION never CAUSATION, HYPOTHESIS never fact, provenance on every measurementIndependent benchmark authority and evaluation standards
Pricing30-day enterprise pilot; 6 commercial packagesPublic benchmarks + private evaluation infrastructure
Best forEnterprises that need to measure and improve how human operators perform with AIEnterprises that need to evaluate and benchmark AI models and applications
Business Model Parallel

Vals AI is the model-side analogue of what MO§ES™ does on the operator side.

The parallel between Vals AI and MO§ES™ is structural. Vals AI combines public benchmark authority — credible, independent rankings that the market trusts — with private evaluation infrastructure that enterprises use to test models in their own contexts. The public benchmarks establish authority. The private infrastructure monetizes that authority through enterprise contracts.

MO§ES™ aims for the same structure on the operator side. Public performative benchmarks — credible, independent rankings of how operators perform — establish authority. Private evaluation infrastructure — bespoke enterprise evals built around an organization's workflows, roles, models, and tasks — monetizes that authority through enterprise contracts.

In the demo dataset, 50 operators produced 1,668 observations across 5 AI providers, 13 benchmark classes, 27 MCP tools, and 12 interventions over a 30-day window. The benchmark classes are the beginning of a public benchmark authority: standardized, repeatable evals that produce comparable results across organizations. The bespoke enterprise evals are the private infrastructure: company-specific evals that answer company-specific questions.

This is why Vals AI is a reference model, not a direct competitor. The two companies operate on different sides of the same structural pattern. Vals AI evaluates the model. MO§ES™ evaluates the operator. Both combine public authority with private infrastructure.

Complementary Evaluation

The model and the operator are both part of the system.

An AI system's output is a function of both the model and the operator. A state-of-the-art model operated poorly produces poor results. A mediocre model operated well can produce good results. Evaluating only the model tells you half the story. Evaluating only the operator tells you the other half.

Vals AI tells you which model to choose. MO§ES™ tells you how well your people use the model you chose. Together they answer the question that matters: how much value is your AI investment actually producing, and where are the bottlenecks — in the model, in the operator, or in the interaction between them?

The interaction is where the most interesting measurement happens. In the demo dataset, 5 AI providers were measured across the same cohort of operators. Some operators performed well across all providers. Some performed well with one provider and poorly with another. This operator-model interaction is invisible to model-only evaluation and to operator-only evaluation. It requires measuring both sides simultaneously — which is what MO§ES™ does when it records which model each operator used.

Reference Model

What the parallel teaches.

Vals AI demonstrates that the public-authority-plus-private-infrastructure model works for AI evaluation. The market values independent benchmarks. Enterprises pay for private evaluation infrastructure. The combination creates a defensible position.

MO§ES™ applies the same pattern to operator evaluation. The public performative benchmarks — standardized metrics like Leverage, Yield, Token SNR, Log Leverage, and Construction — establish authority. The private bespoke evals — enterprise-specific configurations saved as JSON and shareable across CLI, TUI, and MCP — monetize that authority.

The key insight from Vals AI is that benchmark authority is earned through rigor, transparency, and independence. MO§ES™'s governance guardrails — DEVELOPMENTAL gates, ASSOCIATION never CAUSATION, HYPOTHESIS never fact, provenance on every measurement — are designed to earn that same authority. An operator benchmark that claims to rank people without declaring its epistemic status will not be trusted. An operator benchmark that labels every measurement, every diagnosis, and every outcome join with its evidence status can be.

Decision Framework

When to choose which.

Choose Vals AI if
  • You need to evaluate and benchmark AI models
  • You need independent benchmark authority for model selection
  • You want private evaluation infrastructure for your model testing
  • Your question is "which model is best for our use case?"
  • You want to test AI applications before deployment
Choose MO§ES™ if
  • You need to evaluate how human operators perform with AI
  • You need performative benchmarks for operator comparison
  • You want bespoke enterprise evals for your workflows and roles
  • Your question is "how well are our people using AI?"
  • You want to measure, benchmark, diagnose, intervene, and re-evaluate

Most enterprises should use both. Vals AI for the model side. MO§ES™ for the operator side. Together they cover the full AI system.

Related

Go deeper.