Florian Handt · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22811257
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Contemporary AI capability is produced centrally, by training, and delivered frozen at deployment. Downstream, autonomous agents apply that fixed capability through skills — concrete, task-level procedures — whose improvement, even where automated, is governed by static rules: rules that improve skills but do not improve themselves. Two ceilings follow: a volume ceiling (more skills and rules to maintain than any team sustains) and an imagination ceiling (improvement only along axes a human pre-conceived). We call the resulting condition the autonomy paradox — autonomous agents running on a hand-maintained skill base that trends toward maintenance collapse. We present Lineage-Directed Skill Evolution (LDSE), a design theory (Gregor Type V) for a skill layer that improves itself through variation, selection, and inheritance directed by a hard test oracle, in domains where "better" is deterministically decidable. LDSE couples a market (self-organizing selection, curation, retirement) with a reactor (self-improvement through mutation and emergence-gated recombination), and preserves the full provenance of every skill change as a first-class, inspectable artifact. Building on our prior governed-marketplace design theory (AGORA), LDSE treats the hard oracle not as truth but as a fallible referee, and discharges the resulting governance problem — how a self-improving population judges its own improvements without a human in the per-decision loop — with a triad: a fallible selection oracle drives evolution, an independent held-out oracle (never optimized against) detects gaming as an automatic gate, and complete provenance attributes each flagged change to its cause. On this substrate a distinctive product emerges: experience-shaped competence profiles — informally "digital seniors" — and their composition into a Competence Mosaic, which (conditional on the central proposition below) doubles as the decorrelated checker set that keeps the loop honest. The theory's central proposition, Path-Dependent Competence (equal skill + equal knowledge but a different formation path → different generalization), is stated as a testable proposition with a complete, pre-registrable twin-design test protocol; it is the paper's methodological core, not a claimed result. We distinguish our contribution sharply from the growing body of population-based skill/harness evolution (which optimizes against a fitness signal and treats single-lineage search as a pathology to avoid): LDSE is the design theory in which path-dependence is an intended, inspectable, and causally-testable property against a fixed oracle rather than an avoided bug. A pilot instantiation on a 1B-parameter open model supports the selection claim (H1) and demonstrates the mechanism; the causal test of H3 is specified as future work requiring cluster-scale compute.
No comments yet — start the discussion below.