Research paper 02
Metrics Are Not Authority: Control Boundaries for Agentic Evaluation Systems
Agentic software systems now produce logs, scores, traces, harness results, recommendation rows, risk estimates, and routing suggestions at a pace that exceeds ordinary human review cycles. Those measurements are operationally valuable, but they create a recurring control problem: if a system can score an artifact, can it also approve the artifact? If it can detect risk, can it deny action? If it can summarize evidence, can it promote a claim? This paper argues that the answer must be no unless the metric is explicitly delegated authority by a separate governance layer. The central design rule is simple: metrics may describe, warn, compare, rank, route, or recommend, but metrics may not approve, deny, promote, close, revoke, grant capability, or bypass the Controller. The paper develops this rule from Switchboard and Caliper artifacts, including historical Fractal Governance evidence, where evaluation, routing, runtime admission, and publication pipelines are deliberately separated from authority-bearing decisions. It treats metric output as evidence or telemetry, not as permission. The contribution is a control-boundary architecture for agentic evaluation systems: every metric must declare its observation scope, admissible consumers, non-authority status, failure behavior, and claim ceiling. This is not a claim that metrics are unimportant. It is the opposite. Metrics become safer and more useful when their limits are machine-readable.
- Paper
- 02
- Authors
- A.G. Mauro and C.A. Harris
- Date
- 2026-07-19
- Collection
- Standing Framework Research
Abstract
Agentic software systems now produce logs, scores, traces, harness results, recommendation rows, risk estimates, and routing suggestions at a pace that exceeds ordinary human review cycles. Those measurements are operationally valuable, but they create a recurring control problem: if a system can score an artifact, can it also approve the artifact? If it can detect risk, can it deny action? If it can summarize evidence, can it promote a claim? This paper argues that the answer must be no unless the metric is explicitly delegated authority by a separate governance layer. The central design rule is simple: metrics may describe, warn, compare, rank, route, or recommend, but metrics may not approve, deny, promote, close, revoke, grant capability, or bypass the Controller. The paper develops this rule from Switchboard and Caliper artifacts, including historical Fractal Governance evidence, where evaluation, routing, runtime admission, and publication pipelines are deliberately separated from authority-bearing decisions. It treats metric output as evidence or telemetry, not as permission. The contribution is a control-boundary architecture for agentic evaluation systems: every metric must declare its observation scope, admissible consumers, non-authority status, failure behavior, and claim ceiling. This is not a claim that metrics are unimportant. It is the opposite. Metrics become safer and more useful when their limits are machine-readable.
← Back to research papers