Alfredo Madrid-García, Beatriz Merino Barbancho · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23008156
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Background: Jev is a non-generative “System One” model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question- answering and case-based diagnostic-reasoning tasks are unknown. Methods: We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA (1,373 examination questions, 277 of them without a correct substantive option), PubMedQA (500 research questions on abstracts), DiagnosisArena-MCQ (915 published cases) and the NEJM Case Challenges (34 cases with a final diagnosis). GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. Results: All 8,469 requests returned a valid answer. Jev’s accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%; difference 0.2 percentage points, 95% CI −2.2 to 2.6), lower on MetaMedQA (74.8% vs 82.7%;−7.9) and much lower on DiagnosisArena- MCQ (59.8% vs 82.4%;−22.6) and the NEJM cases (61.8% vs 82.4%;−20.6). On MetaMedQA, Jev’s probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability ≥0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev’s probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was “I don’t know or cannot answer”, Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27–0.31 s; all 2,823 items cost US$0.08. Conclusions: Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Its better calibration on examination questions gave no selective-prediction advantage, and it seldom chose “I don’t know” when that was the correct answer. Task-specific validation is required before clinical use.
No comments yet — start the discussion below.