Jolyon Grace · Figshare 2026 · 2026
DOI: 10.6084/m9.figshare.34052538
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Agentic coding and operations work is usually delegated to a model running an open tool loop: the model decides each next step, re-reads its growing transcript every turn, and stops when it believes it has finished. That arrangement is affordable on a provider with prompt caching, where each turn pays only for new tokens. On a stateless provider — most non-frontier APIs, and every small model running on local hardware — the same loop re-reads the full transcript every turn, so cost grows roughly with the square of the number of turns, and the model is never forced to produce an answer. In our production deployment one such run on a stateless model aborted at turn 43 after consuming about 156,000 input tokens to produce 4,300 output tokens, and every stateless agent had to be fenced out of agentic work.This paper reports the design and measured behaviour of Progressive Detail, an implementation — not a new standard — of the Progressive Skill Ontology Standard (PSOS) [2] and of the continuity requirements of the Local Multi-Agent Coordination Protocol (LMCP) [1]. Progressive Detail has four parts: (1) a four-tier description of every retrievable item (trigger, description, summary, detail) with the rule that a model never receives detail unless a cheaper tier selected it; (2) deterministic, LLM-free progressive recall over those tiers; (3) emulated between-turn caching that gives a stateless model the input shape a cached model would see; and (4) a spliced pipeline in which the harness, not the model, drives the work: orient, plan, a bounded number of short single-purpose steps that must each return structured output, review, and synthesis. The harness owns truth: it re-checks every file-and-line citation against the repository, greps for anything the model did not examine, accepts done only when a grounded finding supports it, and refuses role deliverables (a test run, a video, a triaged ticket) that were not actually produced.Measured in the reference deployment, the spliced pipeline turned a simulated looping model from 80 turns and about 253,000 tokens with no verdict into 17 calls and about 27,000 tokens with findings from every step; turned a storyboard task that had consumed 1.33 million input tokens over six failed runs into a single delivered run of 29 calls and about 15,500 input tokens; and let a 3-billion-parameter on-device model answer an exact call-site question with nine atomic calls of at most about 443 tokens each. We report the failures as carefully as the successes: in the first week of production use, fewer than half of stateless runs completed, in about half the cases because of provider credentials, rate limits and gate refusals rather than model behaviour.
No comments yet — start the discussion below.