GC Pranil, Ravinder-Jeet Singh, Ratvinder Grewal · Algorithms 2026 · 2026
DOI: 10.3390/a19090736
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Machine learning models offer high discriminatory power for cross-sectional stroke-status classification, yet high-performing models often remain “black boxes” with little intrinsic interpretability. Bayesian networks (BNs) are probabilistic graphical models that offer intrinsic, auditable interpretability without the need for post hoc explanation. We specify a conditional Gaussian Bayesian network over eleven clinical variables, learned by constrained hill-climbing under epidemiologically derived arc constraints, parameterized by Bayesian estimation for discrete nodes and linear Gaussian regression for continuous ones, and queried for posterior stroke probability by Monte Carlo likelihood weighting. Using a 2 × 2 × 2 factorial design on 5109 patients, we crossed three preprocessing decisions—discretized versus continuous topology, median versus single imputation by chained equations (SICE), and no balancing versus synthetic oversampling (SMOTE)—yielding eight configurations, each assessed for discrimination, for calibration against a prevalence-only reference, and across 30 repeated stratified resamples of the entire pipeline analyzed by linear mixed model. Topology was the decisive choice: continuous nodes raised AUC by 0.050 (95% CI 0.038–0.063, p = 2.5 × 10−9), whereas imputation choice had no detectable effect (p = 0.741). Synthetic oversampling degraded discrimination and interacted antagonistically with topology, harming continuous networks five to six times more than discretized ones (p < 0.001). No configuration exceeded the prevalence-only reference in Brier skill score; a prior correction largely repaired the calibration intercept (−3.01 to −0.04) but not the slope or discrimination loss. The selected network (test AUC 0.811, 95% CI 0.761–0.853) showed no detected difference from logistic regression (AUC 0.819, p = 0.46) or XGBoost (AUC 0.822, p = 0.31). In conclusion, intrinsically interpretable Bayesian networks carried no measured accuracy penalty, though probability estimates require recalibration and longitudinal cohort validation.
No comments yet — start the discussion below.