Research paper 23
Harness Search Under Governance: Optimizing Agent Scaffolds Without Losing Proof Boundaries
Agent performance depends on more than model weights. Context, prompts, tool exposure, memory, retrieval, environment snapshots, retry policy, approval policy, and evidence capture can all change what an agent accomplishes. That makes harness search attractive: instead of treating the scaffold around an agent as fixed, a system can search over candidate scaffold changes. But harness search is dangerous when the search process can mutate the proof boundary. A candidate may improve a score by changing timeout policy, exposing a forbidden tool, weakening evidence requirements, reading development traces, or moving open-track evidence into closed-track claims. This paper proposes harness search under governance: a protocol for optimizing agent scaffolds while freezing the authority, proof, leakage, model, adapter, timeout, environment, and claim boundaries that make comparison meaningful. The operating object follows from Switchboard Harness DSL: Runtime Lab Scenario Packets as Governed Agent Work Contracts by A.G. Mauro and C.A. Harris: the search should run over declared scenario packets and candidate mutation records, not loose prompt experiments. The source evidence is local and design-bounded: Switchboard's Meta-Harness horizon retargets broad search into one native quality-improvement packet, the benchmark contract separates closed-track proof from open-track comparison, and the Terminal-Bench failure-forensics packet shows why opaque failures must be diagnosed before widening. The contribution is a control-plane method, not a leaderboard claim.
- Paper
- 23
- Authors
- A.G. Mauro and C.A. Harris
- Date
- 2026-07-19
- Collection
- Standing Framework Research
Abstract
Agent performance depends on more than model weights. Context, prompts, tool exposure, memory, retrieval, environment snapshots, retry policy, approval policy, and evidence capture can all change what an agent accomplishes. That makes harness search attractive: instead of treating the scaffold around an agent as fixed, a system can search over candidate scaffold changes. But harness search is dangerous when the search process can mutate the proof boundary. A candidate may improve a score by changing timeout policy, exposing a forbidden tool, weakening evidence requirements, reading development traces, or moving open-track evidence into closed-track claims. This paper proposes harness search under governance: a protocol for optimizing agent scaffolds while freezing the authority, proof, leakage, model, adapter, timeout, environment, and claim boundaries that make comparison meaningful. The operating object follows from Switchboard Harness DSL: Runtime Lab Scenario Packets as Governed Agent Work Contracts by A.G. Mauro and C.A. Harris: the search should run over declared scenario packets and candidate mutation records, not loose prompt experiments. The source evidence is local and design-bounded: Switchboard's Meta-Harness horizon retargets broad search into one native quality-improvement packet, the benchmark contract separates closed-track proof from open-track comparison, and the Terminal-Bench failure-forensics packet shows why opaque failures must be diagnosed before widening. The contribution is a control-plane method, not a leaderboard claim.
← Back to research papers