Elizah K Abhinav Victor Korati · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23160777
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
"Preprint. Submitted to IEEE Access (under review)." Autonomous financial dispute-resolution agents that rely on general-purpose, API-hosted large language models (LLMs) inherit two structural limitations: the base model’s weights are inaccessible for adaptation, and its behavior is not shaped by the organization’s own resolved-case history. This paper presents a parameter-efficient fine-tuning (PEFT) approach for inducing a domain-specialized “Auditor Agent” within a deterministic-LLM hybrid orchestration system, extending our prior work on graph-scoped multi-agent chargeback dispute resolution (Sentinel Recover). We formalize the transition from an API-hosted general model to a locally-adapted domain model via Quantized Low-Rank Adaptation (QLoRA), using supervised fine-tuning data harvested directly from production audit trails written back to a relational store during normal agent operation. We introduce a formal verification-checkpoint extension to the underlying finite-state machine S = (V, E, Σ, δ, v0, F ) that governs agent orchestration, and prove that this extension preserves the bounded-termination property of the original system under substitution of the underlying model. We detail the complete data pipeline, including schema design, deduplication, class-balance correction, and prompt-template alignment; the quantization and adapter configuration, including a hyperparameter search space suited to commodity and free-tier GPU compute; the training procedure; and a full evaluation protocol with primary and secondary metrics, an ablation design, and a statistical testing plan for the escalation-frequency comparison against the existing Groq-hosted baseline. We report the architecture, formal properties, and evaluation methodology in full. Consistent with open and honest reporting practice, we do not present placeholder or illustrative empirical numbers in place of measured results: the QLoRA training run against production-scale data has not yet been executed at the time of writing, and a companion technical report will present escalation-frequency, fallback-distribution, verdict-agreement, and latency results once that run and its evaluation are complete.
No comments yet — start the discussion below.