Standing Framework

Research paper 28

Local Model Admission as Governed Infrastructure

Local model infrastructure is often evaluated with simple performance questions: does the model answer, how fast is it, and how many tokens per second can it produce? For governed agent systems, those questions are too narrow. A local model route is not only a provider daemon. It is an admission surface with host fingerprint, provider record, model inventory, startup and warmup behavior, queue wait, load duration, time to first token, total duration, throughput, fallback, cancellation, refusal, and runtime visibility. This paper studies local-model admission as governed infrastructure using Switchboard's local-model benchmark contract. The contract states that the local-model control plane is no longer the open question; the remaining proof burden is whether managed local inference on Apple Silicon is fast, stable, and truthful enough to widen. The contribution is an admission protocol: a model/provider pair becomes useful evidence only when it passes through Switchboard's governed route path, archives enough context for later audit, and remains revocable when host, provider, model, or route conditions drift. The claim ceiling is final-local methods design. Local success, if later observed, would still be limited by host class, provider version, model inventory, route policy, and acceptance thresholds.

Paper
28
Authors
A.G. Mauro and C.A. Harris
Date
2026-07-19
Collection
Standing Framework Research

Abstract

Local model infrastructure is often evaluated with simple performance questions: does the model answer, how fast is it, and how many tokens per second can it produce? For governed agent systems, those questions are too narrow. A local model route is not only a provider daemon. It is an admission surface with host fingerprint, provider record, model inventory, startup and warmup behavior, queue wait, load duration, time to first token, total duration, throughput, fallback, cancellation, refusal, and runtime visibility. This paper studies local-model admission as governed infrastructure using Switchboard's local-model benchmark contract. The contract states that the local-model control plane is no longer the open question; the remaining proof burden is whether managed local inference on Apple Silicon is fast, stable, and truthful enough to widen. The contribution is an admission protocol: a model/provider pair becomes useful evidence only when it passes through Switchboard's governed route path, archives enough context for later audit, and remains revocable when host, provider, model, or route conditions drift. The claim ceiling is final-local methods design. Local success, if later observed, would still be limited by host class, provider version, model inventory, route policy, and acceptance thresholds.

← Back to research papers