Hoda Waguih · Future Business Journal 2026 · 2026
DOI: 10.1186/s43093-026-01017-y
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Machine learning models are increasingly being incorporated into AI-enabled decision-support systems, particularly in high-stakes domains such as healthcare. However, evaluation of such systems continues to rely predominantly on predictive performance metrics, which provide only a partial assessment of whether an AI-based decision-support system can be considered trustworthy for consequential decision-making. Recent advances in trustworthy artificial intelligence (AI) emphasize that characteristics such as explainability, reliability, robustness, and domain relevance should also be considered when evaluating AI systems intended to support decision-making. This study proposes a multidimensional trustworthiness evaluation framework for binary machine learning systems that integrates five dimensions: Predictive Reliability, Explainability, Explanation Stability, Clinical Plausibility, and Imbalance Robustness. The framework is designed as a reusable evaluation methodology rather than as a new classifier, feature-selection method, or explainability algorithm. Its empirical utility was examined through a retrospective case-study evaluation of HER2 status prediction in the METABRIC breast cancer dataset as a representative high-stakes clinical decision-support case. The evaluation included 1980 cases, comprising 1733 HER2-negative and 247 HER2-positive cases, and three classifiers—Decision Tree, Support Vector Machine (SVM), and XGBoost. Models were evaluated using 30 independent repetitions of stratified 80/20 holdout evaluation, with SHAP and LIME used for model explanation. Predictive performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC, while trustworthiness dimensions were operationalized using ROC-AUC, SHAP–LIME agreement, explanation stability, clinical plausibility, and G-mean, respectively. XGBoost achieved the highest ROC-AUC (0.983 ± 0.008), PR-AUC (0.921 ± 0.029), F1-score (0.864 ± 0.024), explanation stability (0.652 ± 0.045), and clinical plausibility (0.303 ± 0.072). Decision Tree achieved the highest SHAP–LIME agreement (0.527 ± 0.105), whereas SVM achieved the highest G-mean (0.926 ± 0.023). Friedman tests identified significant differences among classifiers across all five trustworthiness dimensions ( p < 0.001 for each dimension). Pairwise analyses further demonstrated dimension-specific differences, including significant contrasts between all model pairs for predictive reliability, while Decision Tree and XGBoost did not differ significantly in explainability (Holm-adjusted p = 0.054) or clinical plausibility (Holm-adjusted p = 0.474). These findings demonstrate that model ranking can vary across trustworthiness dimensions and that predictive performance alone does not provide a sufficient characterization of AI systems intended to support consequential decisions. The framework therefore provides a structured multidimensional approach for evaluating binary machine learning systems, with potential relevance to trustworthy AI-enabled decision support beyond the clinical case examined here. The HER2 case study should not be interpreted as evidence of prospective clinical deployment readiness.
No comments yet — start the discussion below.