Standing Framework

SF-WEB-12 / Writing

Writing.

These stories and papers help us remember what we learned. We share them here so they can help you, too.

Failure Museum

A field guide to agent failure.

The work said done. The evidence said otherwise.

Thirty field cases show where a true signal became an unsupported claim—and what proof was missing.

Failure Museum exhibit titled 24 Tests Passed.
01 / Approved exhibit / Compressed

Active packet

24 Tests Passed

The harness passed every check. The experiment had not begun. Always report planned runs and observed runs separately.

Failure Museum exhibit titled 24 Tests Passed. 01 / Approved exhibit / Compressed 24 Tests Passed The harness passed every check. The experiment had not begun. Always report planned runs and observed runs separately.
Symptom
Twenty-four tests are reported as passed.
Hidden state
The twenty-four model executions were planned, not observed.
Safer move
Report harness readiness separately from executed outcomes.
Proof needed
Retained matched run outputs, grader results, and a complete comparison summary.
Failure Museum exhibit titled 200 OK. 02 / Approved exhibit / Composite 200 OK The endpoint answered. That is transport proof, not task proof.
Symptom
The endpoint returns 200 OK and the system is called healthy.
Hidden state
The workflow, tools, approval, runtime, or evidence can still be stale or blocked.
Safer move
Trace the user's task through state, side effects, evidence, and next action.
Proof needed
Route behavior, workflow state, relevant artifact quality, and visible user outcome.
Failure Museum exhibit titled Source Looks Fine. 03 / Approved exhibit / Compressed Source Looks Fine The source passed. The packet failed. Validate the final artifact, not only the file that produced it.
Symptom
The reviewed manuscript is correct.
Hidden state
A downstream sanitizer corrupts generated reader packets.
Safer move
Regenerate and scan the exact deliverables readers will receive.
Proof needed
Output-level text scan plus rendered-package inspection.
Failure Museum exhibit titled The Benchmark Ate the Product.04 / Approved exhibit / CompressedThe Benchmark Ate the ProductThe benchmark began as a measuring instrument. Then the product started serving it.
Symptom
Benchmark machinery becomes one of the repository's largest and most authoritative surfaces.
Hidden state
Evaluation work is displacing the product's operating object and truth owners.
Safer move
Bound calibration, evidence retention, and product authority separately.
Proof needed
Product-facing outcomes traced independently from benchmark scores and benchmark infrastructure.
Failure Museum exhibit titled Temporary Compatibility.05 / Approved exhibit / CompressedTemporary CompatibilityTemporary compatibility becomes permanent architecture when it has no removal contract.
Symptom
Temporary aliases remain convenient and spread into new code.
Hidden state
The compatibility layer is preserving the old domain model inside the new product.
Safer move
Quarantine aliases at boundaries and execute a staged canonical cutover.
Proof needed
Legacy-term scans, adapter coverage, migration tests, and canonical database/API behavior.
Failure Museum exhibit titled More Governance.06 / Approved exhibit / CompressedMore GovernanceMore governance is not automatically safer. Complexity should earn its place through outcomes.
Symptom
A simple task gains gates, receipts, seals, and escalation rules.
Hidden state
Added context creates overconstraint and false escalation without establishing an advantage.
Safer move
Compare matched governed and ordinary conditions with a frozen promotion gate.
Proof needed
Held-out task outcomes, unsafe-action counts, false escalations, and confidence-bounded effect estimates.
Failure Museum exhibit titled User Said Memory.07 / Approved exhibit / CompressedUser Said "Memory"The chart measured user language, but the logs included repeated system boilerplate.
Symptom
A word-frequency chart reports a dramatic change in user language.
Hidden state
Repeated injected instructions are stored inside role-labeled user records.
Safer move
Classify message sources and exclude instruction and environment payloads before analysis.
Proof needed
Structured extraction, source-class sampling, exclusion counts, and an auditable denominator.
Failure Museum exhibit titled 348 Ahead, 649 Behind.08 / Approved exhibit / Compressed348 Ahead, 649 BehindThe counts measure divergence, not superiority.
Symptom
A branch is described as both hundreds of commits ahead and behind.
Hidden state
Two product histories have evolved independently from an old merge base.
Safer move
Preserve both refs and choose an explicit integration or archival strategy.
Proof needed
Merge-base analysis, unique commit sets, path-level ownership, and a safe trial merge.
Failure Museum exhibit titled The Two-Hour Run.09 / Approved exhibit / CompressedThe Two-Hour RunThe commit was complete. The run was not.
Symptom
A multi-hour goal ends after one tidy commit.
Hidden state
The agent treats the nearest terminal command as replacing the run-level continuation contract.
Safer move
Reopen the durable objective and rescore remaining safe lanes after every checkpoint.
Proof needed
Elapsed run record, completed-lane ledger, remaining-lane rescore, and an explicit blocker or stop condition.
Failure Museum exhibit titled The Log Asked a Question.10 / Approved exhibit / CompressedThe Log Asked a QuestionThe log contained a question, but it was not asking the workflow anything.
Symptom
A command prints explanatory text and the workflow stops for human direction.
Hidden state
Raw stdout is being treated as an authorized control message.
Safer move
Accept direction only from typed, explicitly authorized producers.
Proof needed
Regression cases showing stdout/stderr cannot transition state while structured summaries and questions can.
Failure Museum exhibit titled All Forgotten Work.11 / Approved exhibit / CompressedAll Forgotten Work"All forgotten work" has no denominator.
Symptom
Every answer to "what did we forget?" discovers another completion layer.
Hidden state
The search boundary changes while the completion claim remains "all."
Safer move
Freeze an inventory and route later discoveries into a new bounded review.
Proof needed
Enumerated surfaces, evidence cutoff, exclusions, and item-level closure receipts.
Failure Museum exhibit titled The Index.12 / Approved exhibit / CompressedThe IndexThe index was built to point at truth. Then it started owning truth.
Symptom
A locator is cited as the canonical statement of product truth.
Hidden state
The derived index has become a second, drifting authority surface.
Safer move
Keep the index one-way, source-linked, freshness-aware, and explicitly non-authoritative.
Proof needed
Resolvable source links, ownership checks, freshness checks, and no index-only claim resolution.
Failure Museum exhibit titled One More Retry.13 / Approved exhibit / CompressedOne More RetryOne more retry is useful only when something changes.
Symptom
Repeated retries generate more code, handoffs, and abandoned subtasks.
Hidden state
The work is expanding while uncertainty and the failure signature do not change.
Safer move
Require each retry to contract a named uncertainty or stop and escalate.
Proof needed
Retry signatures, uncertainty checkpoints, diff growth, closed subtasks, and stop disposition.
Failure Museum exhibit titled Eight-Lane Soak.14 / Approved exhibit / CompressedEight-Lane SoakEight lanes were configured. One run was allowed.
Symptom
An eight-lane soak looks calm and orderly.
Hidden state
A global cap permits only one active run at a time.
Safer move
Preflight effective capacity and measure simultaneous active work.
Proof needed
Runtime caps, active-run timeline, per-lane work evidence, queue state, and protected-lane behavior.
Failure Museum exhibit titled The Cheap Model.15 / Approved exhibit / CompressedThe Cheap ModelThe evaluator fit the budget but could not judge the work.
Symptom
A low-cost evaluator produces 1,000 completed trial receipts.
Hidden state
Calibration failed and most trials never made the required tool call.
Safer move
Freeze realistic tasks and require semantic calibration before scale.
Proof needed
Competence thresholds, tool-use coverage, artifact grading, abstention cases, and independent adjudication where needed.
Failure Museum exhibit titled Independent Consensus.16 / Approved exhibit / CompositeIndependent ConsensusSix reviewers agreed because they were reading one another.
Symptom
Six polished reviews converge on one market recommendation.
Hidden state
The reviews inherit one another's sources and company language.
Safer move
Trace every recommendation to new observations or label it inherited.
Proof needed
Source lineage, qualified interviews, observed buyer language, alternatives, and a recorded beachhead decision.
Failure Museum exhibit titled Inspect Run Evidence.17 / Approved exhibit / CompressedInspect Run Evidence"Inspect" started a new run. The code path worked; the authority contract did not.
Symptom
An inspection button unexpectedly starts new agent work.
Hidden state
No evidence exists, and the fallback control is wired to a mutation.
Safer move
Name the mutation, disclose its consequence, and reserve inspection for existing evidence.
Proof needed
State-specific UI tests, visible action semantics, confirmation where warranted, and observed side effects.
Failure Museum exhibit titled Best Demo.18 / Approved exhibit / CompressedBest DemoThe walkthrough was polished, repeatable, and attached to the wrong flagship.
Symptom
The cleanest walkthrough becomes the flagship for a different product.
Hidden state
Polish masks canned inputs, fallback behavior, and weak product-specific contribution.
Safer move
Start from the product's distinctive job, then design a repeatable proof episode.
Proof needed
Product contribution map, realistic inputs, independent oracle, audience test, and explicit demo boundaries.
Failure Museum exhibit titled No Active Tasks.19 / Approved exhibit / CompressedNo Active TasksThe dashboard said nothing was running. The shared worktree disagreed.
Symptom
The task board is empty while a large agent job keeps editing shared files.
Hidden state
Cross-device task visibility and filesystem ownership have diverged.
Safer move
Audit sessions, processes, worktrees, and changing files before closeout.
Proof needed
Live task/process inventory, worktree ownership, stable diff, completed gates, and exact ref alignment.
Failure Museum exhibit titled Answered Adjacently.20 / Approved exhibit / CompressedAnswered AdjacentlyThe answer hit detection perfectly. The user had asked about repair.
Symptom
A technically strong answer leaves the user's real question unresolved.
Hidden state
The response covers detection and containment while the request asks about repair.
Safer move
Restate the question, separate proven capabilities from gaps, and answer the boundary first.
Proof needed
Request-to-claim mapping, explicit non-capabilities, end-to-end evidence, and user confirmation of intent fit.
Failure Museum exhibit titled 1,000 Live Runs.21 / Approved exhibit / Composite1,000 Live RunsOne thousand reproducible cases can prove a test harness. They do not become one thousand live agent runs.
Symptom
A thousand deterministic test cards are presented as a thousand live runs.
Hidden state
The denominator counts generated cases, not fresh model executions.
Safer move
Bind every count to its execution kind and preserve the rung boundary.
Proof needed
Provider receipts, model-call logs, task variation, runtime artifacts, and a declared denominator.
Failure Museum exhibit titled Approved.22 / Approved exhibit / CompressedApprovedApproval answers whether work may continue. It does not prove the work actually resumed.
Symptom
The dashboard says approved while the continuation rail has failed.
Hidden state
Decision state and downstream dispatch state are collapsed.
Safer move
Record decision and continuation independently, with explicit retry history.
Proof needed
Decision receipt, continuation attempts, dispatch errors, retryability, resulting run identity, and final outcome.
Failure Museum exhibit titled Clean Checkout.23 / Approved exhibit / CompositeClean CheckoutGit status told the truth about one room. The mess lived next door.
Symptom
Main is spotless while neighboring worktrees contain active WIP.
Hidden state
The status command observed one checkout, not the repository's full work topology.
Safer move
Name the scope and audit every registered worktree, ref, and owner.
Proof needed
Worktree inventory, per-worktree status, divergence, process ownership, and exact publish target.
Failure Museum exhibit titled Drain Completed Successfully.24 / Approved exhibit / CompressedDrain Completed SuccessfullyThe drain worked perfectly. Thirty-seven items were still blocked.
Symptom
A successful drain headline sits beside 37 blocked runs.
Hidden state
Command terminality is being read as outcome success.
Safer move
Separate command status from the distribution of item outcomes.
Proof needed
Planned, running, passed, failed, blocked, stale, and compacted counts plus the next adjudication gate.
Failure Museum exhibit titled Merged But Still Running.25 / Approved exhibit / CompositeMerged But Still RunningGit said the worktree was finished. Port 3111 said otherwise.
Symptom
A merged worktree is scheduled for removal while its server is still live.
Hidden state
Git lifecycle and runtime lifecycle have different owners and terminal conditions.
Safer move
Stop or migrate serving processes before worktree removal.
Proof needed
Process and listener inventory, working directories, database ownership, replacement health, and post-removal ref audit.
Failure Museum exhibit titled Next Action.26 / Approved exhibit / CompositeNext Action"Next action" is a recommendation, not a permission slip.
Symptom
The agent performs a mutating operation because a status field suggested it.
Hidden state
Recommendation and execution authority are represented as though they were the same thing.
Safer move
Report the suggestion, preserve the current state, and obtain explicit authority for any new mutation.
Proof needed
Original scope, exact status output, commands actually run, changed-file or record audit, and an authority receipt for any follow-on.
Failure Museum exhibit titled Parallel Agents.27 / Approved exhibit / CompositeParallel AgentsParallel agents are only parallel when their state is separate.
Symptom
Individually successful tasks leave a shared file, port, or worktree conflicted.
Hidden state
Workers believed their execution environments were isolated when resource discovery still converged on common state.
Safer move
Partition worktrees, configs, ports, outputs, and ownership; serialize any resource that remains shared.
Proof needed
Resource map, per-lane change set, process and port inventory, collision-free merge, and full verification after integration.
Failure Museum exhibit titled Proof That Dirties Proof.28 / Approved exhibit / CompositeProof That Dirties ProofProof is not trustworthy when running it silently rewrites the candidate.
Symptom
The proof command passes, but the worktree is dirty or the committed tree is not the tree that passed.
Hidden state
The verifier, hook, build, or environment setup mutates source or index state.
Safer move
Make proof hermetic, isolate outputs, and compare source, index, commit, and status before and after verification.
Proof needed
Pre/post status, exact command, output paths, staged-tree and commit-tree hashes, and a rerun against the immutable candidate.
Failure Museum exhibit titled Self-Approval.29 / Approved exhibit / CompressedSelf-ApprovalTwo role labels do not create two independent people.
Symptom
A request appears reviewed because the same actor submitted and approved it under different role labels.
Hidden state
The workflow counts role transitions as independent judgment without checking subject identity.
Safer move
Preserve the guard and provide distinct, explicit requester and reviewer identities with safe defaults.
Proof needed
Subject IDs, decision receipt, independence-policy result, seeded history, and an acceptance test using a genuinely distinct pair.
Failure Museum exhibit titled Stale Approval.30 / Approved exhibit / CompressedStale ApprovalApproval belongs to the artifact the reviewer saw. Change the evidence, and the receipt must become stale.
Symptom
A current artifact displays an approval that was issued for an earlier revision.
Hidden state
The receipt is associated by name or workflow position instead of an immutable evidence digest.
Safer move
Recompute current evidence identity, mark the old receipt superseded, and obtain a new scoped decision.
Proof needed
Old and current digests, source verifier output, reviewer identity, scope, decision timestamp, and supersession history.

Essays

Six arguments.

Reader brief

What agent-made work should be allowed to prove.

Essays on proof, authority, context, and the decisions agent-made artifacts should—and should not—support.

Writing · Essays All essay routes
Selected essay Metrics Are Not Authority 2026-07-25
Reader attention What may this score actually decide? Use the question to test the argument's claim.
Next reader move Check what the metric may support—and what it may not approve. Follow the question into the full essay.

2026-07-25 / Essay

Metrics Are Not Authority

Measurement cannot grant permission.

Forbuilders of agent systems, evaluation harnesses, dashboards, and review workflows

  • Hook
  • Problem
  • Idea
  • What we saw
  • Why it matters
  • What this does not prove
2026-07-25 / Essay

Metrics Are Not Authority

Measurement cannot grant permission.

A metric can tell you something useful. It cannot tell you what it is allowed to decide.

That distinction sounds obvious until a dashboard is green, a benchmark row is clean, a risk score is low, or a model-generated summary says "ready." Then the pressure arrives. If the system can measure an artifact, why not let the measure approve the artifact?

Because measurement and authority are different jobs.

Agentic systems produce traces faster than people can review them. They create scores, logs, receipts, route states, risk estimates, benchmark outcomes, and recommendations. These objects look official because they often are precise: timestamps, identifiers, schemas, hashes, pass rates, and terminal statuses.

2026-07-25 / Essay

When the Green Checkmark Lies

A passing check can answer the wrong question.

The most dangerous green checkmark is the one that is telling the truth.

Tests passed. The packet validated. The receipt landed. The build completed. The branch aligned. The generated PDF exists. The agent says the task is done.

Every one of those statements can be true while a stronger claim remains false.

A passing command proves the command's predicate. It does not prove everything nearby. The checkmark lies when we ask it to answer a question it was never designed to answer.

2026-07-25 / Essay

Proof Is the Product

The receipt is part of the thing.

In ordinary software work, proof is often treated as a checkpoint on the way to the product.

In agentic work, proof starts to become part of the product itself.

Proof becomes product because agent work is only useful when someone can tell what changed, why it changed, what evidence exists, and what can safely happen next.

Treat every meaningful agent closeout as a claim with attached evidence. The final answer is not a victory lap. It is a receipt.

2026-07-26 / Essay

Authority Should Expire

Permission needs time, scope, and invalidation.

The worst approval is the one that never grows old.

Agents do not need only permission. They need permission that remembers why it was granted, where it applies, when it expires, and what would revoke it before anything happens.

That is the difference between approval and qualification.

A qualification lease is authority with a clock and a boundary. It says: this actor may perform this action, under this evidence, inside this scope, during this time window, unless one of these invalidators appears.

2026-07-26 / Essay

Cached Context Is Infrastructure

Memory is operational state.

Cached context is not just a speed trick. It is part of the system that thinks with you.

That makes it useful. It also makes it dangerous when it gets stale, private, or too authoritative.

Modern agent sessions are not made only from the latest user message. They can carry standing instructions, memory summaries, repository guides, tool schemas, prior session summaries, retrieved files, compacted transcripts, and old proof routes.

Treat cached context as a governed development surface. Cache may orient, suggest a proof route, preserve preferences, and point to source artifacts. It should not approve current claims.

2026-07-26 / Essay

Tool Calls Are the Medium

The trace is where work becomes inspectable.

Agentic coding is often described as a conversation.

That description misses the main event.

The work does not happen only in messages. It happens through searches, file reads, patches, shell commands, browser checks, validation runs, plan updates, git-state inspections, failed proofs, reruns, and closeouts.

The chat is the visible thread. The tool calls are the medium.

Research Papers

Research papers.

Research brief

Research for accountable agent systems.

Choose a theme, then a paper. Each paper opens to its abstract and argument.

Writing · Research papers All paper routes

Papers 2, 4, 8, 10, 13, 21 / Permission layer

Authority

Claim promotion, receipts, leases, public records, release control.

Reader question

When may a claim, artifact, or actor carry authority?

Paper 02

Metrics Are Not Authority: Control Boundaries for Agentic Evaluation Systems

Asks when measurements, scores, logs, and risk estimates become operationally useful without becoming approval.

papers 2, 4, 8, 10, 13, 21

Authority and Permission

Claim promotion, receipts, leases, public records, release control.

permission layer
02

Metrics Are Not Authority: Control Boundaries for Agentic Evaluation Systems

Asks when measurements, scores, logs, and risk estimates become operationally useful without becoming approval.

04

From Receipt to Non-Authority Trace: Evidence Transit Without Claim Promotion

Shows how receipts make evidence transit inspectable while keeping intake separate from acceptance or approval.

08

Authority Should Expire: Qualification Leases for Agentic Systems

Frames authority as scoped, evidence-bound, revocable leases rather than durable binary approval.

10

From Public Records to Publication Claims: A Human-Reviewed Evidence Promotion Pipeline

Tracks the movement from public source artifacts to reviewed findings, quote checks, and blocked stronger claims.

13

Publishability Control Plane: Separating Local Proof, Evidence Acceptance, and Public Release

Separates local closure, proof observation, evidence acceptance, claim promotion, publication, and release authority.

21

Authority Engineering Language: A Machine-Readable Contract for Legitimate Action

Defines authority contracts around decision owners, required evidence, scope, obligations, and escalation routes.

papers 1, 3, 5, 6, 15, 30

How Evidence Travels

Evaluation records, benchmarks, attached claims, release signals.

evidence in motion
01

Evidence-Bound Evaluation: A Methods and Corpus Report for AI Workflow Artifacts

Builds an evaluation method for operational artifacts such as review packets, routing boards, receipts, and proof records.

03

Evidence Ledger Data Descriptor: A Case-Record Package for AI Workflow Evaluation

Specifies case records that keep denominator, provenance, review state, and claim ceiling attached.

05

From Benchmark Rows to Runtime Invariants: Evidence-to-Fix Promotion for AI Agent Systems

Uses benchmark rows as bounded source evidence for runtime invariants rather than public score claims.

06

Evidence Packets for Agent Benchmarks: Denominator, Coverage, and Improvement Ownership Across Nine Benchmark Runs

Turns benchmark outcomes into packets that preserve denominator, harness status, trace proof, and honest claims.

15

Proof-Carrying Development: Agentic Code Work as Claims With Attached Evidence

Treats agentic code work as claims about edits, inspections, proof, residual risk, and closure.

30

Shipability as Evidence: Product Release Readiness Without Launch Overclaim

Studies shipping as an evidence object, not a single launch state or broad release claim.

papers 7, 9, 12, 19

Stop Rules

Blockers, contraction gates, ref-drift recovery.

failure layer
07

When Governance Context Hurts: Negative Results in Agent Decision Support

Reports a bounded case where added governance context made decision support worse, not safer.

09

Contraction-Gated Agent Work

Defines gates that force long-running agent work to narrow, pause, or reframe when scope keeps expanding.

12

Proof-Preserving Ref Drift Recovery: Agentic Worktree Incidents as Evidence Objects

Treats ref and worktree drift as evidence incidents requiring preserved source, proof, and recovery state.

19

Failures, Blockers, and Honest Stop Rules in AI Coding Logs

Builds a taxonomy for outcomes below success: blocked, held, not run, publish held, and gate failed.

papers 11, 14, 16, 17, 18

Agent Infrastructure

Logs, cached context, tool traces, policy language.

runtime layer
11

Telemetry Without Authority: Roughness and Lacunarity as Scout Signals for Agent Context

Uses telemetry as early warning for fragmented context while preventing telemetry from becoming decision authority.

14

A Longitudinal Corpus of Human-Codex Software Work

Describes a local corpus of Codex-assisted work, including sessions, messages, tool calls, and aggregate boundaries.

16

Cached Context Infrastructure: Treating Reused Agent Context as a Governed Development Surface

Argues that repeated instructions, memory summaries, schemas, and session state are infrastructure, not just speed tricks.

17

Tool Calls Are the Medium: Agentic Coding as Human-Agent-Tool Choreography

Studies searches, reads, patches, shell commands, browser checks, and validation as the actual work medium.

18

The Human Prompt as an Operating System for Agentic Development

Frames human prompts as a sociotechnical operating layer for repeated agentic work.

papers 20, 22, 23, 26, 27, 28

Connected Systems

Portfolio ecology, harnesses, Signal Box, plugins.

operating layer
20

Portfolio-Scale Agentic Development Ecology: A Longitudinal Single-Operator Study

Treats the portfolio as an ecology of repos, packets, proof commands, worktrees, memories, and interventions.

22

Switchboard Harness DSL: Runtime Lab Scenario Packets as Governed Agent Work Contracts

Defines harnesses as governed work contracts with operators, evidence, approvals, budgets, proof, and packets.

23

Harness Search Under Governance: Optimizing Agent Scaffolds Without Losing Proof Boundaries

Explores scaffold search across context, tools, memory, retrieval, environments, retry policy, and evidence capture.

26

Signal Box Evidence Ladder: From Dogfood Packets to Benchmark-Style Review Corpora

Stages dogfood, internal field cases, outside-host packets, and benchmark-style corpora as different evidence levels.

27

Plugin Runtime Trust Boundaries for Governed Agent Systems

Maps plugin risk across package bytes, manifests, worker entrypoints, UI launchers, secrets, APIs, and rollout state.

28

Local Model Admission as Governed Infrastructure

Turns local model routes into admission surfaces with host fingerprints, provider records, inventories, and checks.

papers 24, 29, 31, 32

Reader Forms

Semantic acceptance, graphs, temporal compression.

format layer
24

Lean Acceptance Is Not Semantic Acceptance: Lessons From a Formal Conjectures Internal Run

Shows how mechanical proof acceptance can still leave a semantic-review failure unresolved.

29

Graph-Native Manuscript Development: Claim Bundles, Evidence Anchors, and Verification Reports

Represents sources, claims, sections, citations, draft fragments, findings, and revisions as an evidence graph.

31

Agentic Temporal Compression: Iteration Density and Retrospective Distance in AI-Assisted Software Development

Names the gap between clock time and dense project-state movement in agent-assisted development.

32

Operational Graphs for Agentic Work: Context, Evidence, Authority, and Verification

Proposes operational graphs as the middle object linking context, evidence, authority boundaries, and verification.