SAKIB MOSTAKIM BHUIYAN · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22925608
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
A common working assumption in privacy auditing is that a well-tuned shadow-model adversary represents something close to a worst-case estimate of membership-inference risk, and that cheaper diagnostics such as loss-threshold tests are conservative stand-ins for it. We test this assumption directly on five tabular classification benchmarks (Breast Cancer, Wine, Digits, German Credit, and Adult) using two model families (random forests of varying depth and L2-regularized logistic regression of varying strength), ten random seeds, and three training-set sizes. Across 900 independently trained target models, a calibrated Shokri-style shadow-model attack does not outperform a simple loss-threshold attack on four of the five datasets, and an offline likelihood-ratio attack (LiRA) consistently beats both. We then examine the Confidence-Gap Privacy Score (CGPS) — the difference in mean true-class confidence between training and held-out records — as a cheap, threshold-free proxy that a practitioner could compute without training any shadow models. Pooled across configurations, CGPS correlates strongly with attack AUC (Pearson r between 0.46 and 0.93 depending on dataset and attack), and the association survives partial correlation controlling for the raw loss gap. However, stratifying by model family reveals that most of the pooled signal is carried by random forests (r up to 0.97); for logistic regression the relationship is markedly weaker and, for the shadow and LiRA attacks on the Adult dataset, statistically indistinguishable from zero (r = -0.17, 95% CI [-0.37, 0.01]). We argue that CGPS is best understood as a useful, model-family-dependent screening heuristic rather than a general-purpose substitute for attack simulation, and we report family-stratified rather than pooled statistics as the primary result.
No comments yet — start the discussion below.