Standing Framework

Research paper 16

Cached Context Infrastructure: Treating Reused Agent Context as a Governed Development Surface

Large-context agent workflows increasingly depend on cached context: repeated system instructions, memory summaries, repository guidance, tool schemas, prior transcript slices, and session state. Cache is often treated as a performance detail. This paper argues that cached context is development infrastructure and should be governed as such. The A Longitudinal Corpus of Human-Codex Software Work aggregate recomputation at snapshot cutoff 2026-07-21T07:19:35.322Z reported 20,704,900,411 input tokens, of which 19,684,191,488 were cached input tokens. These figures are local and private. They are not billing evidence, productivity evidence, or a model-quality result, but they show that cached input dominated the observed token surface of long-running agent work. The contribution is a cached-context control model: cache entries need provenance, freshness, authority, invalidation rules, privacy handling, and claim boundaries. The goal is not to reduce cache use. It is to prevent stale or over-authoritative cached material from silently steering code, research, or publication decisions. Cached context should accelerate work without becoming hidden authority.

Paper
16
Authors
A.G. Mauro and C.A. Harris
Date
2026-07-21
Collection
Standing Framework Research

Abstract

Large-context agent workflows increasingly depend on cached context: repeated system instructions, memory summaries, repository guidance, tool schemas, prior transcript slices, and session state. Cache is often treated as a performance detail. This paper argues that cached context is development infrastructure and should be governed as such.

The A Longitudinal Corpus of Human-Codex Software Work aggregate recomputation at snapshot cutoff 2026-07-21T07:19:35.322Z reported 20,704,900,411 input tokens, of which 19,684,191,488 were cached input tokens. These figures are local and private. They are not billing evidence, productivity evidence, or a model-quality result, but they show that cached input dominated the observed token surface of long-running agent work.

The contribution is a cached-context control model: cache entries need provenance, freshness, authority, invalidation rules, privacy handling, and claim boundaries. The goal is not to reduce cache use. It is to prevent stale or over-authoritative cached material from silently steering code, research, or publication decisions. Cached context should accelerate work without becoming hidden authority.

← Back to research papers