Brian C. Long · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23128391
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Coding-agent evaluations typically score one agent on one task in a clean sandbox. Operating a fleet of agents against a live repository raises a different question: what happens when agents produce changes faster than people can understand them? This working paper proposes two constructs for that question. Delegation debt is the accumulated gap between the authority, consequence, and volume of decisions delegated to autonomous systems and the organization's capacity to audit, reconstruct, control, recover from, and defend those decisions. The latency of legibility is the time between an agent action and the moment a qualified reviewer can say what changed, why, on what evidence, and with what downstream effect. We give a simple stipulated model for the first, standard queueing results for the second, and an information-theoretic statement of what a fluent agent-written summary can and cannot establish as evidence. We illustrate the setting with an experience report from the author's own agent fleet, including summary statistics computed from a dated export of 3,209 open pull requests and two days of measured throughput of the repository's final verification step, and we propose design requirements and an evaluation plan. This version reports no controlled experiment: the model is conceptual, the fleet statistics describe volume rather than quality, and the design requirements are proposals that have not been tested. Contents of this record: the compiled PDF (version 0.2.1, 20 pages) and a source bundle with the LaTeX source, bibliography, figure scripts, the aggregate series behind Figure 1 (snapshot counts, submission and merge series, gate-run tuples), CITATION.cff, and the pre-registration draft, fixed-seed sampling script, verified-fact summary generator with its regression tests, and power simulation for the proposed seeded-defect pilot. No pull-request-level dataset is released. Fleet statistics describe volume and throughput, not quality; no human review was in the path of the submitted changes; the gate whose outage is reported is the author's company's own software (see the paper's Declarations).
No comments yet — start the discussion below.