Dawa Chyophel Lepcha, Aaliya Ali, Sophie Martin, Shabbir Syed-Abdul · bioRxiv (Cold Spring Harbor Laboratory) 2026 · 2026
DOI: 10.64898/2026.08.15.744687
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Explainability methods applied to deep learning models for Alzheimer’s disease neuroimaging produce attribution maps that vary substantially across methods and architectures, yet no validated quantitative framework exists for determining which method most faithfully localises attribution signal within established AD neuroimaging biomarker anatomy at the individual subject level. Existing validation approaches rely on group-level comparisons against meta-analytic activation maps or qualitative visual inspection, leaving individual-level biomarker alignment uncharacterised across the full cognitively normal, mild cognitive impairment, and AD diagnostic spectrum. We introduce the Biomarker Fidelity Score (BFS), a quantitative clinical AI validation tool that measures the spatial overlap between individual-level three-dimensional explainability attention maps and atlas-registered AD-relevant neuroimaging biomarker regions of interest across thirteen anatomically defined structures including the hippocampus, entorhinal cortex, amygdala, and parahippocampal gyrus. Five established explainability methods (GradCAM++, Integrated Gradients, DeepSHAP, Layer-wise Relevance Propagation, and ScoreCAM) were benchmarked across three volumetric architectures (3D ResNet-18, DenseNet-121, and Swin-UNETR) on 327 balanced ADNI-3 subjects. Integrated Gradients achieved the highest BFS across all three architectures while GradCAM++ consistently showed the lowest biomarker alignment (all p<0.001, Friedman test). The complete BFS pipeline, applied without retraining, replicated these method rankings with near-perfect fidelity on an independent cohort of 207 OASIS-3 subjects, with a maximum absolute difference of 0.0005 across all fifteen method-architecture combinations and a Spearman rank correlation of 0.964 between cohort rankings, providing rare direct evidence that XAI method reliability generalises across scanners, acquisition protocols, and populations. By offering an externally validated, individual-level, biomarker-grounded quantitative standard, the BFS framework equips clinicians and AI developers with practical translational guidance for selecting trustworthy explainability methods in AD neuroimaging, directly supporting the responsible clinical deployment of explainable AI.
No comments yet — start the discussion below.