Nicolás Vera Zúñiga · arXiv (Cornell University) 2026 · 2026
DOI: 10.5281/zenodo.23038894
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Code for the paper Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores. A language model reads a clinical note into facts marked present, absent or unknown, each with an exact evidence quote; deterministic code computes the range of scores still possible given the unknowns and asks the clinician only about inputs that could change the decision category. Six calculators (HEART, CURB-65, qSOFA, PERC, Wells, Cockcroft-Gault), a seeded synthetic cohort generator, note rendering and validation, extraction, five question-asking policies, a simulated clinician, evaluation and the scripts that regenerate every table and figure of the paper. Accompanies arXiv:2609.34112 (cs.CL). Only code is released. The synthetic notes, extractions and model-response cache are not included; the cohort is regenerated from the seed, and re-running the pipeline re-queries the models. MedCalc-Bench Verified, used as a real-note check, is not redistributed.
No comments yet — start the discussion below.