Concept · Governance

AI operator evaluation governance.

Governance in AI operator evaluation defines the evidence labels, decision-use constraints, and privacy boundaries that govern how measurements can be used. The core principle is simple: results route workflows, not people. No punitive use. No automatic adverse employment actions.

DEVELOPMENTAL HYPOTHESIS ASSOCIATION VALIDATION REQUIRED
Evidence Labels

Every measurement carries its epistemic status.

The system never collapses measurement, diagnosis, intervention, and validation into a single claim. Every output carries an evidence label that declares what kind of claim it is — and what it is not.

MEASURED SIGNAL ≠ DERIVED METRIC ≠ BENCHMARK ≠ HYPOTHESIS ≠ VALIDATED OUTCOME
PROVEN

Established through prior validation. The strongest evidence label. Rare in the current framework — most claims are not yet independently validated.

MEASURED

Direct observation from telemetry. Token counts, session counts, timestamps. The raw data surface.

OBSERVED

Direct observation of structural signals. Provider fields, model identifiers, tool usage patterns.

DERIVED

Computed from telemetry via declared formula. All five canonical metrics are DERIVED. Not independently validated as performance measures unless separately tested.

HYPOTHESIS

Diagnosis from pattern analysis. Carries evidence, alternatives, and HYPOTHESIS status. Never presented as established fact.

PILOT

Results from the pilot program. Based on synthetic data by default. Not yet validated against real-world outcomes.

Decision-Use Labels

How results can and cannot be used.

Decision-use labels are enforced in code. They determine what actions a measurement can inform — and what actions it cannot. These are not guidelines. They are constraints built into the system.

DEVELOPMENTAL

Production gate results route workflows, not people. DEVELOPMENTAL measurements can inform workflow design, tooling decisions, training programs, and process improvements. They cannot be used for adverse employment actions — hiring, firing, promotion, compensation, or performance review.

ASSOCIATION

Outcome joins between interventions and business metrics are labeled ASSOCIATION — never CAUSATION. A correlation between an intervention and a business metric is not proof that the intervention caused the business change. Internal metric deltas and external outcome deltas are kept in separate fields.

VALIDATION REQUIRED

Claims that have not been independently validated carry the VALIDATION REQUIRED label. This includes most derived metrics, most benchmark positions, and all outcome joins. The label is not a warning — it is a declaration of epistemic status.

Core Constraints

What the system will not do.

Governance is defined as much by what the system refuses to do as by what it does. These constraints are non-negotiable in the pilot and carry forward into production.

  • No punitive use. Measurements are not used to punish operators. DEVELOPMENTAL labels mean results route workflows, not people.
  • No automatic adverse employment actions. The system does not generate, recommend, or trigger adverse employment actions — hiring, firing, promotion, compensation, or performance review. This is enforced in code, not just policy.
  • No employee ranking for personnel decisions. Benchmark positions and composite scores are structural signals for workflow optimization, not personnel evaluation tools.
  • No surveillance by default. The system works from token counts and structural signals, not prompt content. Cohort-level reporting is the default. Individual identity requires separate authorization.
  • No collapsing evidence levels. The system never presents a DERIVED metric as a MEASURED observation, a HYPOTHESIS as a fact, or an ASSOCIATION as a CAUSATION. Evidence labels are enforced in code.
  • No retroactive metric selection. Intervention target metrics are declared before the intervention begins. The system does not retroactively select metrics to make results look favorable.
Provenance

Every measurement carries its history.

Provenance is not optional. Every measurement, metric, benchmark, diagnosis, and intervention record carries a full provenance chain that traces it back to its source.

  • Source telemetry window and operator identifiers
  • Evidence label (PROVEN, MEASURED, OBSERVED, DERIVED, HYPOTHESIS)
  • Decision-use label (DEVELOPMENTAL, ASSOCIATION, VALIDATION REQUIRED)
  • Synthetic-data flag (pilot runs on synthetic data by default)
  • Validation status (VALIDATION REQUIRED for unvalidated claims)
  • Intervention target and follow-up window (for intervention records)

Provenance means that any claim made by the system can be traced back to the telemetry it was computed from, the formula that was applied, the cohort it was benchmarked against, and the evidence label it carries. This is not a logging feature — it is a governance requirement.

Privacy Boundaries

Evaluate operation without surveillance.

The system can work from telemetry and structural signals rather than requiring full prompt-content inspection. This is not a claim of absolute privacy — it is a design principle that minimizes data collection.

DATA MINIMIZATION
  • Content-free or content-minimized telemetry where possible
  • Token counts, not prompt text
  • Structural signals, not conversation content
  • Cohort-level reporting by default
  • Individual identity requires separate authorization
GOVERNANCE CONTROLS
  • Configurable retention
  • Enterprise-defined usage policies
  • No adverse employment action in pilot
  • Provenance on every measurement
  • Decision-use labels enforced in code
Falsifiability

The system is designed to be wrong.

If an intervention does not produce a measurable change in the target metric, the system says so. If a diagnosis cannot be distinguished from an alternative, the system says so. If an outcome join cannot be validated, the system says so.

This is a formal research framework, not established science. It carries papers, experiments, datasets, patents, and falsifiability work. It is presented with appropriate epistemic humility.

Related Concepts

Read alongside.

The content-free data foundation that governance protects.

ASSOCIATION labels in action — never CAUSATION.

DEVELOPMENTAL label on the 0–100 index.

Read the Methodology Request a Pilot