Standing Framework

Research paper 23

Harness Search Under Governance: Optimizing Agent Scaffolds Without Losing Proof Boundaries

Agent performance depends on more than model weights. Context, prompts, tool exposure, memory, retrieval, environment snapshots, retry policy, approval policy, and evidence capture can all change what an agent accomplishes. That makes harness search attractive: instead of treating the scaffold around an agent as fixed, a system can search over candidate scaffold changes. But harness search is dangerous when the search process can mutate the proof boundary. A candidate may improve a score by changing timeout policy, exposing a forbidden tool, weakening evidence requirements, reading development traces, or moving open-track evidence into closed-track claims. This paper proposes harness search under governance: a protocol for optimizing agent scaffolds while freezing the authority, proof, leakage, model, adapter, timeout, environment, and claim boundaries that make comparison meaningful. The operating object follows from Switchboard Harness DSL: Runtime Lab Scenario Packets as Governed Agent Work Contracts by A.G. Mauro and C.A. Harris: the search should run over declared scenario packets and candidate mutation records, not loose prompt experiments. The source evidence is local and design-bounded: Switchboard's Meta-Harness horizon retargets broad search into one native quality-improvement packet, the benchmark contract separates closed-track proof from open-track comparison, and the Terminal-Bench failure-forensics packet shows why opaque failures must be diagnosed before widening. The contribution is a control-plane method, not a leaderboard claim.

Paper
23
Authors
A.G. Mauro and C.A. Harris
Date
2026-07-19
Collection
Standing Framework Research

Abstract

Agent performance depends on more than model weights. Context, prompts, tool exposure, memory, retrieval, environment snapshots, retry policy, approval policy, and evidence capture can all change what an agent accomplishes. That makes harness search attractive: instead of treating the scaffold around an agent as fixed, a system can search over candidate scaffold changes. But harness search is dangerous when the search process can mutate the proof boundary. A candidate may improve a score by changing timeout policy, exposing a forbidden tool, weakening evidence requirements, reading development traces, or moving open-track evidence into closed-track claims. This paper proposes harness search under governance: a protocol for optimizing agent scaffolds while freezing the authority, proof, leakage, model, adapter, timeout, environment, and claim boundaries that make comparison meaningful. The operating object follows from Switchboard Harness DSL: Runtime Lab Scenario Packets as Governed Agent Work Contracts by A.G. Mauro and C.A. Harris: the search should run over declared scenario packets and candidate mutation records, not loose prompt experiments. The source evidence is local and design-bounded: Switchboard's Meta-Harness horizon retargets broad search into one native quality-improvement packet, the benchmark contract separates closed-track proof from open-track comparison, and the Terminal-Bench failure-forensics packet shows why opaque failures must be diagnosed before widening. The contribution is a control-plane method, not a leaderboard claim.

← Back to research papers