Research paper 06
Evidence Packets for Agent Benchmarks: Denominator, Coverage, and Improvement Ownership Across Nine Benchmark Runs
Agent benchmark rows often collapse several different questions into one score: did the agent finish the task, did the harness run correctly, did the environment expose the required tools, did the trace preserve proof, and what can anyone honestly claim afterward? This paper argues for evidence packets as the missing bridge between benchmark rows and reviewable runtime records. The source draft describes a local whatdoyouwant benchmark-evidence corpus with nine validated evidence packets spanning 4,189 benchmark rows across terminal-use, tool-use, scientific, biology, machine-learning-engineering, and formal-mathematics tasks. The archived corpus summary reports 4,165 Switchboard proof rows, 13 explicit missing-evidence rows, and improvement ownership split across none 2,358, model 1,213, adapter 573, benchmark_env 32, and evidence 13. These numbers are not a leaderboard. They are a denominator-preserving map of what the runtime could prove, what failed, what was missing, and who or what would need to improve. The central claim is modest: benchmark rows become more useful for governed agent systems when transformed into evidence packets that preserve provenance, denominator state, run proof, failure classes, improvement ownership, and claim ceilings.
- Paper
- 06
- Authors
- A.G. Mauro and C.A. Harris
- Date
- 2026-07-19
- Collection
- Standing Framework Research
Abstract
Agent benchmark rows often collapse several different questions into one score: did the agent finish the task, did the harness run correctly, did the environment expose the required tools, did the trace preserve proof, and what can anyone honestly claim afterward? This paper argues for evidence packets as the missing bridge between benchmark rows and reviewable runtime records. The source draft describes a local whatdoyouwant benchmark-evidence corpus with nine validated evidence packets spanning 4,189 benchmark rows across terminal-use, tool-use, scientific, biology, machine-learning-engineering, and formal-mathematics tasks. The archived corpus summary reports 4,165 Switchboard proof rows, 13 explicit missing-evidence rows, and improvement ownership split across none 2,358, model 1,213, adapter 573, benchmark_env 32, and evidence 13. These numbers are not a leaderboard. They are a denominator-preserving map of what the runtime could prove, what failed, what was missing, and who or what would need to improve. The central claim is modest: benchmark rows become more useful for governed agent systems when transformed into evidence packets that preserve provenance, denominator state, run proof, failure classes, improvement ownership, and claim ceilings.
← Back to research papers