Active packet
24 Tests Passed
The harness passed every check. The experiment had not begun. Always report planned runs and observed runs separately.
SF-WEB-12 / Writing
These stories and papers help us remember what we learned. We share them here so they can help you, too.
Failure Museum
Thirty field cases show where a true signal became an unsupported claim—and what proof was missing.
Active packet
The harness passed every check. The experiment had not begun. Always report planned runs and observed runs separately.
01 / Approved exhibit / Compressed
24 Tests Passed
The harness passed every check. The experiment had not begun. Always report planned runs and observed runs separately.
02 / Approved exhibit / Composite
200 OK
The endpoint answered. That is transport proof, not task proof.
03 / Approved exhibit / Compressed
Source Looks Fine
The source passed. The packet failed. Validate the final artifact, not only the file that produced it.
04 / Approved exhibit / CompressedThe Benchmark Ate the ProductThe benchmark began as a measuring instrument. Then the product started serving it.
05 / Approved exhibit / CompressedTemporary CompatibilityTemporary compatibility becomes permanent architecture when it has no removal contract.
06 / Approved exhibit / CompressedMore GovernanceMore governance is not automatically safer. Complexity should earn its place through outcomes.
07 / Approved exhibit / CompressedUser Said "Memory"The chart measured user language, but the logs included repeated system boilerplate.
08 / Approved exhibit / Compressed348 Ahead, 649 BehindThe counts measure divergence, not superiority.
09 / Approved exhibit / CompressedThe Two-Hour RunThe commit was complete. The run was not.
10 / Approved exhibit / CompressedThe Log Asked a QuestionThe log contained a question, but it was not asking the workflow anything.
11 / Approved exhibit / CompressedAll Forgotten Work"All forgotten work" has no denominator.
12 / Approved exhibit / CompressedThe IndexThe index was built to point at truth. Then it started owning truth.
13 / Approved exhibit / CompressedOne More RetryOne more retry is useful only when something changes.
14 / Approved exhibit / CompressedEight-Lane SoakEight lanes were configured. One run was allowed.
15 / Approved exhibit / CompressedThe Cheap ModelThe evaluator fit the budget but could not judge the work.
16 / Approved exhibit / CompositeIndependent ConsensusSix reviewers agreed because they were reading one another.
17 / Approved exhibit / CompressedInspect Run Evidence"Inspect" started a new run. The code path worked; the authority contract did not.
18 / Approved exhibit / CompressedBest DemoThe walkthrough was polished, repeatable, and attached to the wrong flagship.
19 / Approved exhibit / CompressedNo Active TasksThe dashboard said nothing was running. The shared worktree disagreed.
20 / Approved exhibit / CompressedAnswered AdjacentlyThe answer hit detection perfectly. The user had asked about repair.
21 / Approved exhibit / Composite1,000 Live RunsOne thousand reproducible cases can prove a test harness. They do not become one thousand live agent runs.
22 / Approved exhibit / CompressedApprovedApproval answers whether work may continue. It does not prove the work actually resumed.
23 / Approved exhibit / CompositeClean CheckoutGit status told the truth about one room. The mess lived next door.
24 / Approved exhibit / CompressedDrain Completed SuccessfullyThe drain worked perfectly. Thirty-seven items were still blocked.
25 / Approved exhibit / CompositeMerged But Still RunningGit said the worktree was finished. Port 3111 said otherwise.
26 / Approved exhibit / CompositeNext Action"Next action" is a recommendation, not a permission slip.
27 / Approved exhibit / CompositeParallel AgentsParallel agents are only parallel when their state is separate.
28 / Approved exhibit / CompositeProof That Dirties ProofProof is not trustworthy when running it silently rewrites the candidate.
29 / Approved exhibit / CompressedSelf-ApprovalTwo role labels do not create two independent people.
30 / Approved exhibit / CompressedStale ApprovalApproval belongs to the artifact the reviewer saw. Change the evidence, and the receipt must become stale.Essays
Reader brief
Essays on proof, authority, context, and the decisions agent-made artifacts should—and should not—support.
2026-07-25 / Essay
Measurement cannot grant permission.
Forbuilders of agent systems, evaluation harnesses, dashboards, and review workflows
Measurement cannot grant permission.
A metric can tell you something useful. It cannot tell you what it is allowed to decide.
That distinction sounds obvious until a dashboard is green, a benchmark row is clean, a risk score is low, or a model-generated summary says "ready." Then the pressure arrives. If the system can measure an artifact, why not let the measure approve the artifact?
Because measurement and authority are different jobs.
Agentic systems produce traces faster than people can review them. They create scores, logs, receipts, route states, risk estimates, benchmark outcomes, and recommendations. These objects look official because they often are precise: timestamps, identifiers, schemas, hashes, pass rates, and terminal statuses.
A passing check can answer the wrong question.
The most dangerous green checkmark is the one that is telling the truth.
Tests passed. The packet validated. The receipt landed. The build completed. The branch aligned. The generated PDF exists. The agent says the task is done.
Every one of those statements can be true while a stronger claim remains false.
A passing command proves the command's predicate. It does not prove everything nearby. The checkmark lies when we ask it to answer a question it was never designed to answer.
The receipt is part of the thing.
In ordinary software work, proof is often treated as a checkpoint on the way to the product.
In agentic work, proof starts to become part of the product itself.
Proof becomes product because agent work is only useful when someone can tell what changed, why it changed, what evidence exists, and what can safely happen next.
Treat every meaningful agent closeout as a claim with attached evidence. The final answer is not a victory lap. It is a receipt.
Permission needs time, scope, and invalidation.
The worst approval is the one that never grows old.
Agents do not need only permission. They need permission that remembers why it was granted, where it applies, when it expires, and what would revoke it before anything happens.
That is the difference between approval and qualification.
A qualification lease is authority with a clock and a boundary. It says: this actor may perform this action, under this evidence, inside this scope, during this time window, unless one of these invalidators appears.
Memory is operational state.
Cached context is not just a speed trick. It is part of the system that thinks with you.
That makes it useful. It also makes it dangerous when it gets stale, private, or too authoritative.
Modern agent sessions are not made only from the latest user message. They can carry standing instructions, memory summaries, repository guides, tool schemas, prior session summaries, retrieved files, compacted transcripts, and old proof routes.
Treat cached context as a governed development surface. Cache may orient, suggest a proof route, preserve preferences, and point to source artifacts. It should not approve current claims.
The trace is where work becomes inspectable.
Agentic coding is often described as a conversation.
That description misses the main event.
The work does not happen only in messages. It happens through searches, file reads, patches, shell commands, browser checks, validation runs, plan updates, git-state inspections, failed proofs, reruns, and closeouts.
The chat is the visible thread. The tool calls are the medium.
Research Papers
Research brief
Choose a theme, then a paper. Each paper opens to its abstract and argument.
Paper 02
Asks when measurements, scores, logs, and risk estimates become operationally useful without becoming approval.
Claim promotion, receipts, leases, public records, release control.
Asks when measurements, scores, logs, and risk estimates become operationally useful without becoming approval.
Shows how receipts make evidence transit inspectable while keeping intake separate from acceptance or approval.
Frames authority as scoped, evidence-bound, revocable leases rather than durable binary approval.
Tracks the movement from public source artifacts to reviewed findings, quote checks, and blocked stronger claims.
Separates local closure, proof observation, evidence acceptance, claim promotion, publication, and release authority.
Defines authority contracts around decision owners, required evidence, scope, obligations, and escalation routes.
Evaluation records, benchmarks, attached claims, release signals.
Builds an evaluation method for operational artifacts such as review packets, routing boards, receipts, and proof records.
Specifies case records that keep denominator, provenance, review state, and claim ceiling attached.
Uses benchmark rows as bounded source evidence for runtime invariants rather than public score claims.
Turns benchmark outcomes into packets that preserve denominator, harness status, trace proof, and honest claims.
Treats agentic code work as claims about edits, inspections, proof, residual risk, and closure.
Studies shipping as an evidence object, not a single launch state or broad release claim.
Blockers, contraction gates, ref-drift recovery.
Reports a bounded case where added governance context made decision support worse, not safer.
Defines gates that force long-running agent work to narrow, pause, or reframe when scope keeps expanding.
Treats ref and worktree drift as evidence incidents requiring preserved source, proof, and recovery state.
Builds a taxonomy for outcomes below success: blocked, held, not run, publish held, and gate failed.
Logs, cached context, tool traces, policy language.
Uses telemetry as early warning for fragmented context while preventing telemetry from becoming decision authority.
Describes a local corpus of Codex-assisted work, including sessions, messages, tool calls, and aggregate boundaries.
Argues that repeated instructions, memory summaries, schemas, and session state are infrastructure, not just speed tricks.
Studies searches, reads, patches, shell commands, browser checks, and validation as the actual work medium.
Frames human prompts as a sociotechnical operating layer for repeated agentic work.
Portfolio ecology, harnesses, Signal Box, plugins.
Treats the portfolio as an ecology of repos, packets, proof commands, worktrees, memories, and interventions.
Defines harnesses as governed work contracts with operators, evidence, approvals, budgets, proof, and packets.
Explores scaffold search across context, tools, memory, retrieval, environments, retry policy, and evidence capture.
Stages dogfood, internal field cases, outside-host packets, and benchmark-style corpora as different evidence levels.
Maps plugin risk across package bytes, manifests, worker entrypoints, UI launchers, secrets, APIs, and rollout state.
Turns local model routes into admission surfaces with host fingerprints, provider records, inventories, and checks.
Semantic acceptance, graphs, temporal compression.
Shows how mechanical proof acceptance can still leave a semantic-review failure unresolved.
Represents sources, claims, sections, citations, draft fragments, findings, and revisions as an evidence graph.
Names the gap between clock time and dense project-state movement in agent-assisted development.
Proposes operational graphs as the middle object linking context, evidence, authority boundaries, and verification.