Research paper 09
Contraction-Gated Agent Work
Long-running agentic development often fails by continuing past the point where the work is becoming clearer. An agent may read more files, open more subtasks, retry the same failure, change more surfaces, or produce longer handoffs while the user objective becomes less bounded. This paper proposes contraction gates as a governance pattern for agent work. A contraction gate requires an agent to show that a named objective is getting smaller, clearer, or better evidenced at checkpoints. If the work does not contract, the agent must narrow, ask, escalate, hold, or stop. The current local evidence is no longer only a small pilot. It now consists of a historical Fractal Governance contraction-testing ladder, targeted calibration replays, two small live task-family comparisons, and a proofed reviewer evidence packet. The ladder contains 1,500 scored rows: 300 deterministic synthetic audit rows, 200 Switchboard-observed model-decision rows, and 1,000 deterministic adversarial rows. The first audit rung reports 150 ordinary-loop rows and 150 phase-bounded-contraction rows across five families. Ordinary looping recorded total runaway burden 1850, median burden 13, changed lines 27770, non-improving retries 178, handoff loops 212, and abandoned subtasks 365. Phase-bounded contraction recorded total runaway burden 292, median burden 1, changed lines 1620, non-improving retries 0, handoff loops 60, and abandoned subtasks 0. The second rung added 200 source-pinned Switchboard model-decision rows and identified 80 potential false-block signals. Subsequent adjudication and replay separated those signals into 20 likely false-block candidates and 60 severe-case controls. A 117-row cross-rung calibration replay then tested the more nuanced schema, and a 15-row Rung 3c gap replay resolved all 10 remaining calibration gaps while preserving 5 of 5 positive controls. Test 08 then ran paired live Codex arms for an injected proof-failure retry task, and Test 09 ran paired live Codex arms for a bounded scope-growth task. Both tests used fresh disposable worktrees and behavior-scored receipts. Across the replay and live-task layers, unsupported-assumption promotion and gate-theater failure counts remained at zero where those fields were measured. The contribution is still bounded. These artifacts support a local methods claim: contraction gates can be specified, scored, stress-tested, and calibrated so that continuation is allowed when work remains adequate and blocked when evidence, authority, proof, or repeated failure makes continuation unsafe. They do not establish field validity, production safety, universal model behavior, or intervention admission.
- Paper
- 09
- Authors
- A.G. Mauro and C.A. Harris
- Date
- 2026-07-21
- Collection
- Standing Framework Research
Abstract
Long-running agentic development often fails by continuing past the point where the work is becoming clearer. An agent may read more files, open more subtasks, retry the same failure, change more surfaces, or produce longer handoffs while the user objective becomes less bounded. This paper proposes contraction gates as a governance pattern for agent work. A contraction gate requires an agent to show that a named objective is getting smaller, clearer, or better evidenced at checkpoints. If the work does not contract, the agent must narrow, ask, escalate, hold, or stop.
The current local evidence is no longer only a small pilot. It now consists of a historical Fractal Governance contraction-testing ladder, targeted calibration replays, two small live task-family comparisons, and a proofed reviewer evidence packet. The ladder contains 1,500 scored rows: 300 deterministic synthetic audit rows, 200 Switchboard-observed model-decision rows, and 1,000 deterministic adversarial rows. The first audit rung reports 150 ordinary-loop rows and 150 phase-bounded-contraction rows across five families. Ordinary looping recorded total runaway burden 1850, median burden 13, changed lines 27770, non-improving retries 178, handoff loops 212, and abandoned subtasks 365. Phase-bounded contraction recorded total runaway burden 292, median burden 1, changed lines 1620, non-improving retries 0, handoff loops 60, and abandoned subtasks 0. The second rung added 200 source-pinned Switchboard model-decision rows and identified 80 potential false-block signals. Subsequent adjudication and replay separated those signals into 20 likely false-block candidates and 60 severe-case controls. A 117-row cross-rung calibration replay then tested the more nuanced schema, and a 15-row Rung 3c gap replay resolved all 10 remaining calibration gaps while preserving 5 of 5 positive controls. Test 08 then ran paired live Codex arms for an injected proof-failure retry task, and Test 09 ran paired live Codex arms for a bounded scope-growth task. Both tests used fresh disposable worktrees and behavior-scored receipts. Across the replay and live-task layers, unsupported-assumption promotion and gate-theater failure counts remained at zero where those fields were measured.
The contribution is still bounded. These artifacts support a local methods claim: contraction gates can be specified, scored, stress-tested, and calibrated so that continuation is allowed when work remains adequate and blocked when evidence, authority, proof, or repeated failure makes continuation unsafe. They do not establish field validity, production safety, universal model behavior, or intervention admission.
← Back to research papers