Mohammad Kawsar · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23181098
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Deep learning classifiers for polycystic ovary syndrome (PCOS) diagnosis from transvaginal ultrasound are commonly accompanied by post-hoc explanations such as Grad-CAM, which are typically validated only qualitatively, by visual inspection of a small number of heatmaps. This paper presents a quantitative audit of one such pipeline, combining a data-quality check with a faithfulness evaluation across three test conditions of increasing distributional distance from training data. We find that (1) the public PCOSGen dataset contains substantial internal near-duplication (32.2% of training images share a near-identical counterpart) and cross-split contamination (17.4% of the official test set duplicates a training image), but excluding these duplicates does not meaningfully change reported accuracy; (2) a naive random train/test split does not measurably inflate performance relative to a duplicate-aware split, despite 23.3% of its test images having a near-duplicate in training; (3) Grad-CAM explanations are reliably more informative than a random-saliency baseline under an insertion-AUC metric across all three test conditions (p < 2 × 10⁻³¹ in every case), but are statistically indistinguishable from random under a deletion-AUC metric, an asymmetry that becomes complete (p = 0.84) on an independent external dataset; and (4) classification accuracy collapses to 31.1% — below the majority-class baseline — on an external dataset whose class balance is inverted relative to training, despite an AUC of 0.72 that would suggest reasonable discriminative ability. These findings suggest that a single faithfulness metric, and a single in-distribution accuracy figure, are each insufficient to certify that an explainable clinical imaging model is trustworthy under real-world deployment conditions.
No comments yet — start the discussion below.