
Jonas Henn, Alisa Stoll, Philipp Feodorovici, D Subramani, Johannes Röttgen, Jan Arensmeyer, Ingo Gräff, Benjamin Wulff, Jörg C. Kalff, Hanno Matthaei · Scientific Reports 2026 · 2026
DOI: 10.1038/s41598-026-73738-7
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Clinical information required for surgical data science (SDS) is frequently embedded in unstructured text. We developed and evaluated a reproducible pipeline for selecting locally deployed open-weight large language models (LLMs) for binary symptom annotation. In this retrospective single-center study, 1,100 German emergency-department reports were manually annotated for nausea, vomiting, diarrhea, and dysuria. After reserving 100 reports for prompt formulation and temperature testing, nine LLMs were screened on symptom-specific stratified development sets ( N = 250). Selected models were compared with a negation-aware rule-based baseline in independent validation sets ( N = 750) using F 1 -score and patient-level bootstrap confidence intervals. Temperature 0.0 provided the greatest overall stability. Validation F 1 -scores were 0.985 for vomiting, 0.979 for nausea, 0.824 for dysuria, and 0.814 for diarrhea. Corresponding baseline F 1 -scores were 0.913, 0.724, 0.705, and 0.853, respectively. Paired comparisons favored LLMs for nausea and vomiting; confidence intervals included zero for diarrhea and dysuria. Median inference times ranged from 0.298 to 1.653 s per report. Discrepancies reflected operational criteria, temporal variation, inconsistent documentation, missed mentions, and five reference errors. Pragmatic model screening can identify suitable local LLMs for clinical free-text annotation. The pipeline is reproducible and adaptable but requires context-specific configuration and validation.
No comments yet — start the discussion below.