Gonzalo Esteban Mosquera Rojas, Sebastian R. van der Voort, Carolin M. Pirkl, Sandeep S Kaushik, Marion Smits, Stefan Klein · The Journal of Machine Learning for Biomedical Imaging 2026 · 2026
DOI: 10.59275/j.melba.2026-456d
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ within a multi-task Deep Learning (DL) framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts Isocitrate Dehydrogenase (IDH) mutation status, 1p/19q co-deletion status and tumor grade. We use Monte Carlo Dropout (MCD) as the primary sampling-based UQ method for the detailed task-aware analysis, obtaining predictive uncertainty and its aleatoric and epistemic components. We assess uncertainty along complementary axes: (i) MC sample convergence of uncertainty estimates and their decomposition into aleatoric and epistemic components, (ii) calibration of predictive probabilities, and (iii) operational utility of uncertainty estimates, including error detection, selective prediction, and associations with tumor segmentation performance. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE) to assess whether the observed operational utility of uncertainty estimates extends beyond a single UQ method. For the tumor segmentation task, we further examine how different voxel-wise uncertainty aggregation strategies influence case-level reliability assessment, thereby explicitly accounting for task-specific uncertainty representation. Additionally, we study task interactions to quantify how tumor segmentation quality and uncertainty relate to the prediction of the tumor features. Finally, we explore whether a composite trust score integrating tumor segmentation and classification uncertainty improves error detection. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable predictive probabilities. Decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. The comparison with DE and MCDE showed that ensemble-based uncertainty estimates provided comparable operational utility, although no UQ method consistently dominated across all tasks and metrics. Moreover, the proposed trust score did not consistently outperform classification uncertainty for selective prediction, indicating that task-specific predictive uncertainty remains the most informative operational indicator of trust. Overall, our results provide a task-aware evaluation strategy and practical guidance towards the development of trustworthy AI for glioma diagnosis.
No comments yet — start the discussion below.