Standing Framework

Research paper 34

A Replication Corpus for Human-Codex Software Work

This paper applies the aggregate-only method from A Longitudinal Corpus of Human-Codex Software Work to a second single-account Codex JSONL history. The snapshot cutoff is 2026-07-25T19:32:16.892Z. The parser enumerates 13,277 session files: 13,268 included active-root files and 9 included archived-root files. Of those, 12,084 sessions contain usable token counters and 1,193 do not. The first observed event timestamp is 2026-02-10T15:14:13.129Z; the latest included event timestamp is 2026-07-25T19:32:16.748Z. Summing the last cumulative token counter per token-instrumented session yields 47,043,309,796 total tokens, including 46,861,245,371 input tokens, 44,925,306,624 cached input tokens, 180,514,025 output tokens, and 63,263,061 reasoning output tokens. The same parse counts 38,983 user messages, 225,294 assistant messages, 539,831 function calls, and 523,942 shell command calls. The contribution is a replication-style aggregate descriptor that supports Paper 14's method claims across a second private Codex-log corpus without turning either corpus into a public transcript release, billing record, productivity measure, model comparison, or population-general workflow claim.

Paper
34
Authors
Date
2026-07-25
Collection
Standing Framework Research

Abstract

This paper applies the aggregate-only method from A Longitudinal Corpus of Human-Codex Software Work to a second single-account Codex JSONL history. The snapshot cutoff is 2026-07-25T19:32:16.892Z. The parser enumerates 13,277 session files: 13,268 included active-root files and 9 included archived-root files. Of those, 12,084 sessions contain usable token counters and 1,193 do not. The first observed event timestamp is 2026-02-10T15:14:13.129Z; the latest included event timestamp is 2026-07-25T19:32:16.748Z. Summing the last cumulative token counter per token-instrumented session yields 47,043,309,796 total tokens, including 46,861,245,371 input tokens, 44,925,306,624 cached input tokens, 180,514,025 output tokens, and 63,263,061 reasoning output tokens. The same parse counts 38,983 user messages, 225,294 assistant messages, 539,831 function calls, and 523,942 shell command calls. The contribution is a replication-style aggregate descriptor that supports Paper 14's method claims across a second private Codex-log corpus without turning either corpus into a public transcript release, billing record, productivity measure, model comparison, or population-general workflow claim.

← Back to research papers