Yoshiyasu Takefuji · Biomedical Signal Processing and Control 2026 · 2026
DOI: 10.1016/j.bspc.2026.111577
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Supervised machine learning models are increasingly applied in biomedical signal processing and control for feature assessment; however, a critical epistemological gap persists in that no established ground truth exists against which feature importance can be validated. This paper theoretically establishes that supervised models exhibit two distinct dimensions of accuracy, target prediction accuracy and feature importance accuracy, which are not interchangeable. High target prediction accuracy does not guarantee reliable feature importance, as importance metrics reflect contributions to prediction rather than true causal associations. Furthermore, SHAP-based explanations inherit and may amplify the underlying model’s distortions rather than correct them, rendering their outcomes contingent on model-specific artifacts rather than confirmed causal insights. To address this gap, we draw on the extended Bradford Hill criteria to propose consistency and dose–response relationships as principled validation criteria, and we introduce a novel leave-top-feature-out diagnostic framework to empirically test the stability of feature importance rankings. Empirical analysis of a late-life depression dataset comparing nine feature selection algorithms demonstrates that unsupervised models yield superior stability in feature ranking, whereas supervised models exhibit label-driven instability that SHAP integration fails to resolve, underscoring the need for validation frameworks that extend beyond model-dependent interpretability methods.
No comments yet — start the discussion below.