Standing Framework

Research paper 06

Evidence Packets for Agent Benchmarks: Denominator, Coverage, and Improvement Ownership Across Nine Benchmark Runs

Agent benchmark rows often collapse several different questions into one score: did the agent finish the task, did the harness run correctly, did the environment expose the required tools, did the trace preserve proof, and what can anyone honestly claim afterward? This paper argues for evidence packets as the missing bridge between benchmark rows and reviewable runtime records. The source draft describes a local whatdoyouwant benchmark-evidence corpus with nine validated evidence packets spanning 4,189 benchmark rows across terminal-use, tool-use, scientific, biology, machine-learning-engineering, and formal-mathematics tasks. The archived corpus summary reports 4,165 Switchboard proof rows, 13 explicit missing-evidence rows, and improvement ownership split across none 2,358, model 1,213, adapter 573, benchmark_env 32, and evidence 13. These numbers are not a leaderboard. They are a denominator-preserving map of what the runtime could prove, what failed, what was missing, and who or what would need to improve. The central claim is modest: benchmark rows become more useful for governed agent systems when transformed into evidence packets that preserve provenance, denominator state, run proof, failure classes, improvement ownership, and claim ceilings.

Paper
06
Authors
A.G. Mauro and C.A. Harris
Date
2026-07-19
Collection
Standing Framework Research

Abstract

Agent benchmark rows often collapse several different questions into one score: did the agent finish the task, did the harness run correctly, did the environment expose the required tools, did the trace preserve proof, and what can anyone honestly claim afterward? This paper argues for evidence packets as the missing bridge between benchmark rows and reviewable runtime records. The source draft describes a local whatdoyouwant benchmark-evidence corpus with nine validated evidence packets spanning 4,189 benchmark rows across terminal-use, tool-use, scientific, biology, machine-learning-engineering, and formal-mathematics tasks. The archived corpus summary reports 4,165 Switchboard proof rows, 13 explicit missing-evidence rows, and improvement ownership split across none 2,358, model 1,213, adapter 573, benchmark_env 32, and evidence 13. These numbers are not a leaderboard. They are a denominator-preserving map of what the runtime could prove, what failed, what was missing, and who or what would need to improve. The central claim is modest: benchmark rows become more useful for governed agent systems when transformed into evidence packets that preserve provenance, denominator state, run proof, failure classes, improvement ownership, and claim ceilings.

← Back to research papers