Jorge A Carrasco · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23106385
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large language models (LLMs) frequently produce plausible yet incorrect outputs when tasked with specialized domains such as legal document analysis or software engineering. While fine-tuning and retrieval-augmented generation (RAG) are common remedies, they introduce significant computational overhead, architectural complexity, or privacy risks when proprietary knowledge is involved. We introduce a Governed Inference Framework that reduces domain-specific errors by injecting expert-governed context at inference time without modifying model weights. The framework orchestrates task characterization, expert selection, bounded response generation, and independent verification. We report a controlled ablation study across five experimental conditions (A–E) and five paired repetitions (350 total outputs) on a frozen synthetic bank of 14 specialized cases. A 26B-parameter model operating within the full framework (Condition E) achieved 71.4% exact responses (50/70) and mean F1 = 0.932, compared to 0% exact (0/70, F1 = 0.179) for a professional prompt (Condition A), 21.4% exact (15/70, F1 ≈ 0.50) for static expert context (Condition B), and 8.6% exact (6/70, F1 = 0.436) for orchestration without verification (Condition D). Critically, Conditions B and C received identical context bytes (SHA-256 verified), yet Condition C (NeuroSkill runtime) preserved operational capabilities validation, modular selection, versioning, budget control, and verifiable receipts that Condition B lacked. A paired sign-test confirmed that the full framework outperformed the prompt baseline in all 14 cases (p = 0.000122) and outperformed orchestration without verification in 13 of 14 cases (p = 0.000244). We complement the semantic study with an independent operational confirmatory suite (56/56 tests passed) demonstrating that the full system enforces fail-closed behavior: an independent evaluator detects incorrect outputs, and a publication gate raises a Document Quality Error before any non-exact result reaches the user or codebase. The operational matrix assigns 85/100 to the NeuroSkill runtime, 95/100 to the orchestrator, and 100/100 to the complete system. These results support a systems-level thesis: sustained accuracy in specialized domains does not arise solely from the model or from context, but from the complete operational and verification envelope. We identify four systematic failure modes that remain unresolved, report the results as preliminary, and discuss implications for local deployment of open-weight models in regulated industries.
No comments yet — start the discussion below.