Julio Suárez-Albanchez · International Journal of Artificial Intelligence Tools 2026 · 2026
DOI: 10.1142/s0218213026500211
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Artificial Intelligence (AI) systems are increasingly deployed in high-impact domains such as healthcare, cybersecurity, autonomous systems, and industrial decision-making, where reliable model evaluation is essential. Despite continuous improvements in predictive performance, growing evidence indicates that many reported gains are inflated by methodological flaws, including data leakage, shortcut learning, dataset shift, and insufficient reproducibility practices. These issues compromise model generalization, reduce robustness under real-world conditions, and undermine trust in AI systems. This study proposes a reproducible, leakage-aware evaluation framework designed to detect and quantify methodological risks throughout the machine learning pipeline. The framework integrates four complementary dimensions: dataset integrity analysis, pipeline validation, robustness assessment under controlled perturbations, and reproducibility evaluation. These components are consolidated into a novel Leakage Risk Score (LRS), a normalized metric that estimates methodological risk and complements conventional performance measures such as accuracy. A proof-of-concept evaluation was conducted using multiple supervised classification datasets under both clean and intentionally contaminated conditions. The results demonstrate that data leakage consistently inflates predictive performance while reducing robustness and reproducibility, leading to misleading estimates of model quality. Although the empirical validation is intentionally limited and illustrative, the proposed framework establishes a unified methodological foundation for leakage-aware evaluation and offers a practical tool for improving transparency, reproducibility, and reliability in machine learning research. The approach supports the broader transition from benchmark-oriented assessment toward risk-aware evaluation of trustworthy and sustainable AI systems. Largescale empirical validation is left for future work.
No comments yet — start the discussion below.