Attila Kővári · Computers 2026 · 2026
DOI: 10.3390/computers15090560
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Stable discrimination can coexist with unreliable probabilities, abstention policies, and uncertainty sets under distribution shift. We evaluate this mismatch with CRIT-AID, an executable multi-domain reliability-audit framework that separates model fitting, probability calibration, operating-rule calibration, and testing and then transports source-derived rules unchanged to target domains. Across four public tabular domains, stable discrimination did not imply stable probability quality, selective operating points, or conformal uncertainty. On identical ACS 2024 records, changing the income target definition left AUROC nearly unchanged, while ECE differed by 0.083; prevalence-intercept alignment reduced this difference to −0.004, showing that target semantics can alter probability reliability without materially changing ranking. Across 27 primary 90% conformal conditions, label-conditional calibration improved worst-class coverage in 18 but worsened it in 9 and usually increased prediction-set size. LightGBM sensitivity changed absolute discrimination without removing the mismatch among reliability dimensions. These findings show that discrimination alone is insufficient evidence for reliable AI decision support: audits should test the transportability of probability mappings, operating rules, class-specific validity, and uncertainty informativeness when deployment conditions or target meanings change.
No comments yet — start the discussion below.