Research paper 33
Evidence-Bound Harness Engineering: Building Agent Systems Where Proof Travels
Evidence-bound harness engineering is the discipline of designing agent harnesses so work can move quickly while proof, authority, and claim boundaries remain inspectable. That is the category this paper names. Ordinary harness engineering asks how to make an agent more capable: better context, better tools, better memory, better runners, better feedback. Evidence-bound harness engineering asks a second question: when the agent produces something, what is that result allowed to mean? That second question changes the harness. A harness is no longer just the prompt, the tool list, the memory layer, the test runner, or the eval suite. Those things still matter, but they are not enough. Once an agent is doing consequential work, the harness also has to preserve what work was attempted, what evidence was produced, what proof ran, what authority applied, what failed, what stayed in the denominator, and what can honestly be claimed afterward. Without that structure, a fluent completion note can start to sound like done work. A green check can start to sound like release approval. A benchmark row can start to sound like model truth. A receipt can start to sound like accepted evidence. An evidence-bound harness keeps those meanings from drifting. It lets agents move quickly, but it makes the proof travel with the work. That is what makes the category worth naming: not more ceremony around agents, but a way to keep speed from turning into proof debt.
- Paper
- 33
- Authors
- A.G. Mauro and C.A. Harris
- Date
- 2026-09-02
- Collection
- Standing Framework Research
Abstract
Evidence-bound harness engineering is the discipline of designing agent harnesses so work can move quickly while proof, authority, and claim boundaries remain inspectable.
That is the category this paper names. Ordinary harness engineering asks how to make an agent more capable: better context, better tools, better memory, better runners, better feedback. Evidence-bound harness engineering asks a second question: when the agent produces something, what is that result allowed to mean?
That second question changes the harness. A harness is no longer just the prompt, the tool list, the memory layer, the test runner, or the eval suite. Those things still matter, but they are not enough. Once an agent is doing consequential work, the harness also has to preserve what work was attempted, what evidence was produced, what proof ran, what authority applied, what failed, what stayed in the denominator, and what can honestly be claimed afterward. Without that structure, a fluent completion note can start to sound like done work. A green check can start to sound like release approval. A benchmark row can start to sound like model truth. A receipt can start to sound like accepted evidence.
An evidence-bound harness keeps those meanings from drifting. It lets agents move quickly, but it makes the proof travel with the work. That is what makes the category worth naming: not more ceremony around agents, but a way to keep speed from turning into proof debt.
← Back to research papers