A number can warn me. It cannot approve the work unless the system says exactly who gave it that job.
I have learned to distrust any metric that arrives at the meeting already wearing a little badge.
Not because numbers are useless. I love a good count the way I love a straight fence line: it saves argument, exposes neglect, and lets tired people stop guessing. A clean denominator can calm a room. A sharp score can point to the broken hinge before the whole gate drags in the dirt. A dashboard, when it is honest about what it saw and what it missed, can keep a team from driving around the same pasture all afternoon pretending memory is a system.
But I have also watched a metric do that strange startup magic trick where it begins the week as an observation and ends the week as permission. Nobody votes on the promotion. Nobody says, out loud, "the score is now the authority." The sentence just changes shape. A scan found no blocker, therefore the artifact is ready. A benchmark row passed, therefore the system is better. A classifier did not complain, therefore the risk must be gone. A receipt says the packet landed, therefore the evidence must have been accepted.
That is how a useful instrument becomes a small, well-formatted tyrant.
The temptation is real. I do not want to pretend otherwise. When the launch deck is due, the support queue is making little throat-clearing noises in the corner, and the founder brain starts whispering that one more approval ceremony will kill momentum, a green number looks like mercy. It looks like the tired person's way out. It looks like the system being helpful for once.
Sometimes it is helpful. That is the problem.
A metric can be correct inside its own fence and still be wrong for the decision being made. It can observe local execution without seeing customer value. It can validate schema without approving a claim. It can summarize a bounded run without granting release authority. It can flag context risk without deciding that the user intent is invalid. The metric is not lying in those cases. The humans, agents, dashboards, or closeout prose are spending it at the wrong store.
I keep coming back to this because agent systems are not short on traces. They produce logs, scores, receipts, route states, risk estimates, benchmark outcomes, recommendations, hashes, terminal statuses, and little green phrases that smell like closure if you do not stand too close. Precision gives these objects a costume of authority. A timestamp looks official. A schema looks adult. A pass rate looks like it knows where the bodies are buried.
It does not. It knows what it was instrumented to know.
So the rule I want is plain: let metrics do honest metric work. Let them describe, warn, compare, rank, route, and recommend. Let them wake up the reviewer, flag the stale packet, point at the suspicious denominator, and make the next responsible action harder to ignore. Do not let them approve, deny, promote, close, revoke, grant capability, bless a release, launder a publication claim, or sneak past an authority gate unless a separate authority object explicitly delegates that power.
That separate object matters. I want the system to name who can approve, what they can approve, for which artifact, over which source state, with which proof, inside which scope, and under what revocation or expiration conditions. Call it a Controller decision, a review record, a lease, an approval packet, or whatever fits the machinery. I care less about the label than the separation. The number should not be allowed to wake up one morning as sheriff, judge, lender, launch committee, and auntie at the graduation party who knows exactly who has been acting foolish.
This is where a lot of dashboards go soft. They show the green thing and hide the boundary. They say "passed" and make the viewer go hunting for the predicate that passed. They say "ready" when they mean "no blocker found in this scan." They say "valid" when they mean "schema valid." They say "accepted" when they mean "received." That slippage is not cosmetic. It is the path by which a measurement turns into an authority transition without anyone having to admit they moved the gate.
If I were building the surface tomorrow, I would make the metric carry its own humility. Not a vague disclaimer nobody reads, but machine-readable humility: observation scope, denominator, missingness, consumer contract, non-authority clause, failure behavior, and claim ceiling. If the metric is missing, stale, contradictory, out of scope, or non-comparable, the dashboard should say so in the same room where it shows the score. Missingness should not be a footnote stuffed in the glove box. It should stand next to the number with its hat on.
I would also stop letting green be a personality. Green can mean "this bounded predicate passed." It cannot mean "ship it" unless the authority record says shipping is the predicate. Green can mean "the scout saw no smoke from this ridge." It cannot mean "the whole valley is safe." Anybody who has lived around weather, software, or family logistics knows the difference between a useful signal and permission to stop paying attention.
The research basis for this post is narrower than the mood of the rant, and that is important. The source paper comes from local agentic research and development systems where evaluation, routing, runtime admission, publication pipelines, receipts, proofs, and review decisions are deliberately separated. It does not prove that every metric in the world is untrustworthy. It does not prove a general theory of law, compliance, institutional governance, or moral authority. It does not prove that approvals must always be manual. Automated authority can exist, but if the system gives a metric that power, it should have the decency to name the delegation, scope, revocation path, failure behavior, and blast radius before the number starts spending permission it never earned.
My operating rule is simple enough to survive a bad meeting: metrics are scouts, not sheriffs.
Let them ride ahead. Let them notice trouble. Let them come back dusty with a useful report. Then make the authority object decide what happens next.
Source note: based on Metrics Are Not Authority: Control Boundaries for Agentic Evaluation Systems by A.G. Mauro and C.A. Harris.
Claim boundary: this is a public-explainer blog draft for a local methods pattern. It does not approve publication, validate any specific metric, grant legal/compliance authority, prove product quality, or approve public release.
Source trail: related archive work includes From Receipt to Non-Authority Trace, Evidence-Bound Evaluation, Lean Acceptance Is Not Semantic Acceptance, and Shipability as Evidence.