Rodrigo Pagliusi, Leandro G.M. Alvim, Raúl Ferreira, Ygor Canalli, Filipe Braida, Geraldo Zimbrão · International Journal of Data Science and Analytics 2026 · 2026
DOI: 10.1007/s41060-026-01294-4
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Machine learning models are susceptible to reproducing and amplifying structural inequities present in historical data, leading to unjust outcomes in high-stakes domains such as criminal justice, healthcare, employment, and financial services. This paper investigates the relationship between predictive performance and algorithmic fairness in widely used classification methods under systematically increased, group-conditional bias in training labels. We introduce Systematic Label Flipping for Fairness Stress Testing (SLF–FST), a reproducible evaluation framework that injects controllable bias into training data and tracks the joint evolution of accuracy and group-based fairness metrics as bias intensity increases. Using Decision Tree, Random Forest, Logistic Regression, and feedforward Neural Network classifiers, we assess robustness across standard benchmark datasets. Our results reveal comparable degradation trends across classifiers; however, Logistic Regression exhibits the most pronounced decline in both predictive accuracy and fairness on the COMPAS dataset, whereas Random Forests, as an ensemble method, demonstrate the greatest robustness to injected bias. Neural Networks and Decision Trees showed intermediate behavior. These findings suggest that SLF–FST can function as a practical pre-deployment auditing mechanism, enabling practitioners to identify failure thresholds and quantify trade-offs between fairness and predictive performance in risk-sensitive systems.
No comments yet — start the discussion below.