Abstract
Agent benchmarks are usually reported as scores. This paper studies a different use: benchmark rows as evidence for runtime invariants. Across Switchboard Runtime Lab work, rows from Terminal-Bench and AgentDojo were not promoted into public score claims. Instead, they were treated as bounded source evidence for hypotheses about the runtime substrate itself. Four cases show the pattern. Blocked lifecycle states must not be reported as ordinary solver failures. A Line must not complete when current authoritative acceptance evidence failed. Declared materialization contracts must gate pre-execution and pre-submit completion. Side-effect-required work must block before adapter invocation when the selected operator lacks the required capability. The contribution is an evidence-to-fix promotion protocol that lets benchmark traces improve an agent runtime without laundering benchmark measurements into illegitimate authority.
← Back to research papers