Standing Framework

Research paper 01

Evidence-Bound Evaluation: A Methods and Corpus Report for AI Workflow Artifacts

Evaluation of AI workflow systems often compresses work into scalar outcomes: pass rate, preference, model score, cost, speed, or adoption. That compression is poorly matched to agentic workflow products whose outputs are operational artifacts: review packets, routing boards, scenario runs, evidence receipts, proof records, and human-review decisions. This paper presents an evidence-bound evaluation method for such artifacts. The unit of evaluation is a case_record whose operating object, source bundle, protocol, produced artifact, verifier, adjudication state, provenance, outcome, and claim ceiling remain inspectable after the run. In the current validated corpus, three product lanes produce 193 records: Signal Box contributes 161 change-packet records, including 143 accepted and 18 failed records; Interlock contributes 17 accepted route-board records; and Test Stand contributes 15 accepted scenario-run records. A live validation on 2026-07-19 confirmed the same totals: 193 total, 175 accepted, and 18 failed. The contribution is not a product-quality claim, benchmark claim, adoption claim, or authority claim. It is a repeatable method for making workflow evidence usable while keeping failed runs, held receipts, blocked gates, role qualification, and claim ceilings in the denominator.

Paper
01
Authors
A.G. Mauro and C.A. Harris
Date
2026-07-19
Collection
Standing Framework Research

Abstract

Evaluation of AI workflow systems often compresses work into scalar outcomes: pass rate, preference, model score, cost, speed, or adoption. That compression is poorly matched to agentic workflow products whose outputs are operational artifacts: review packets, routing boards, scenario runs, evidence receipts, proof records, and human-review decisions. This paper presents an evidence-bound evaluation method for such artifacts. The unit of evaluation is a case_record whose operating object, source bundle, protocol, produced artifact, verifier, adjudication state, provenance, outcome, and claim ceiling remain inspectable after the run. In the current validated corpus, three product lanes produce 193 records: Signal Box contributes 161 change-packet records, including 143 accepted and 18 failed records; Interlock contributes 17 accepted route-board records; and Test Stand contributes 15 accepted scenario-run records. A live validation on 2026-07-19 confirmed the same totals: 193 total, 175 accepted, and 18 failed. The contribution is not a product-quality claim, benchmark claim, adoption claim, or authority claim. It is a repeatable method for making workflow evidence usable while keeping failed runs, held receipts, blocked gates, role qualification, and claim ceilings in the denominator.

← Back to research papers