Standing Framework

Research paper 05

From Benchmark Rows to Runtime Invariants: Evidence-to-Fix Promotion for AI Agent Systems

Agent benchmarks are usually reported as scores. This paper studies a different use: benchmark rows as evidence for runtime invariants. Across Switchboard Runtime Lab work, rows from Terminal-Bench and AgentDojo were not promoted into public score claims. Instead, they were treated as bounded source evidence for hypotheses about the runtime substrate itself. Four cases show the pattern. Blocked lifecycle states must not be reported as ordinary solver failures. A Line must not complete when current authoritative acceptance evidence failed. Declared materialization contracts must gate pre-execution and pre-submit completion. Side-effect-required work must block before adapter invocation when the selected operator lacks the required capability. The contribution is an evidence-to-fix promotion protocol that lets benchmark traces improve an agent runtime without laundering benchmark measurements into illegitimate authority.

Paper
05
Authors
A.G. Mauro and C.A. Harris
Date
2026-07-19
Collection
Standing Framework Research

Abstract

Agent benchmarks are usually reported as scores. This paper studies a different use: benchmark rows as evidence for runtime invariants. Across Switchboard Runtime Lab work, rows from Terminal-Bench and AgentDojo were not promoted into public score claims. Instead, they were treated as bounded source evidence for hypotheses about the runtime substrate itself. Four cases show the pattern. Blocked lifecycle states must not be reported as ordinary solver failures. A Line must not complete when current authoritative acceptance evidence failed. Declared materialization contracts must gate pre-execution and pre-submit completion. Side-effect-required work must block before adapter invocation when the selected operator lacks the required capability. The contribution is an evidence-to-fix promotion protocol that lets benchmark traces improve an agent runtime without laundering benchmark measurements into illegitimate authority.

← Back to research papers