Karman, Hunter · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.17657737
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
v6 (October 2026): Adds Section 7, correcting the training-procedure record of earlier versions: the evaluated Qwen 2.5 adapters are best-validation-loss checkpoints (checkpoint-200, epoch ≈1.38), not final two-epoch checkpoints; the models trained on two corpus snapshots (1,153 and 1,213 rows) rather than one; the validation set was not held out; the corpus contains ~50% exact duplication (595 unique prompt–response pairs). Section 7 also reports Wilson intervals, restates the evaluation as 57 judgments over 25 unique prompts, names the one adverse judge pass (Qwen 0.5B / Claude Opus 4, 45.6%), and discloses an uncontrolled judge-label asymmetry. Primary results and reporting conventions are unchanged from v5; 25 of 26 canonical judge passes favor the adapted models. This report studies small-data domain adaptation: fine-tuning language models on a hand-authored corpus of 1,213 examples in continental philosophy and speculative reasoning, created through iterative discussions with LLMs (the model used as an authoring tool, with the researcher directing content and reasoning patterns). Holding the dataset fixed, we fine-tune four base models (Qwen 2.5 0.5B/3B/7B and Llama 3.2 3B) and evaluate each with blind, position-randomized A/B testing against its base model, judged by independent LLMs across three laboratories (Anthropic, OpenAI, Google) and two model generations (2025 and 2026). In-domain win rate rises with base size. On the 2026 frontier-judge panel the fine-tunes reach 86.3% (Qwen 7B), 78.9% (Qwen 3B), and 64.5% (Llama 3B), while out-of-domain performance stays at parity (~50%; no catastrophic forgetting). The 7B is the strongest and most temporally stable. Training cost was a few dollars of cloud GPU per model. v5 (June 2026): adds Qwen 2.5 7B (largest model) and a scaling analysis; integrates a win-rate reconciliation. Earlier versions headlined a 91.2% win rate for the Qwen 3B fine-tune; that was a two-judge pooled value under the original split-subset protocol. Under the uniform multi-judge protocol used throughout, the corresponding 2025 figure is an 88.9% three-judge mean, and 91.2% is retained only as the GPT-4o / Gemini cross-laboratory agreement. The PDF's "Note on Revisions" (§0) gives the full reconciliation. Code, evaluation framework, and representative data samples are released.
No comments yet — start the discussion below.