Research paper 16
Cached Context Infrastructure: Treating Reused Agent Context as a Governed Development Surface
Large-context agent workflows increasingly depend on cached context: repeated system instructions, memory summaries, repository guidance, tool schemas, prior transcript slices, and session state. Cache is often treated as a performance detail. This paper argues that cached context is development infrastructure and should be governed as such. The A Longitudinal Corpus of Human-Codex Software Work aggregate recomputation at snapshot cutoff 2026-07-21T07:19:35.322Z reported 20,704,900,411 input tokens, of which 19,684,191,488 were cached input tokens. These figures are local and private. They are not billing evidence, productivity evidence, or a model-quality result, but they show that cached input dominated the observed token surface of long-running agent work. The contribution is a cached-context control model: cache entries need provenance, freshness, authority, invalidation rules, privacy handling, and claim boundaries. The goal is not to reduce cache use. It is to prevent stale or over-authoritative cached material from silently steering code, research, or publication decisions. Cached context should accelerate work without becoming hidden authority.
- Paper
- 16
- Authors
- A.G. Mauro and C.A. Harris
- Date
- 2026-07-21
- Collection
- Standing Framework Research
Abstract
Large-context agent workflows increasingly depend on cached context: repeated system instructions, memory summaries, repository guidance, tool schemas, prior transcript slices, and session state. Cache is often treated as a performance detail. This paper argues that cached context is development infrastructure and should be governed as such.
The A Longitudinal Corpus of Human-Codex Software Work aggregate recomputation at snapshot cutoff 2026-07-21T07:19:35.322Z reported 20,704,900,411 input tokens, of which 19,684,191,488 were cached input tokens. These figures are local and private. They are not billing evidence, productivity evidence, or a model-quality result, but they show that cached input dominated the observed token surface of long-running agent work.
The contribution is a cached-context control model: cache entries need provenance, freshness, authority, invalidation rules, privacy handling, and claim boundaries. The goal is not to reduce cache use. It is to prevent stale or over-authoritative cached material from silently steering code, research, or publication decisions. Cached context should accelerate work without becoming hidden authority.
← Back to research papers