
Nusirat Ojuolape Gold · Discover Artificial Intelligence 2026 · 2026
DOI: 10.1007/s44163-026-02226-8
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Empirical studies have evaluated overall fairness and reliability of AI grading systems at aggregate levels, while investigation on their consistency remains scarce. This study investigated systematic discrepancy between AI and human grades across varying conditions. It further examined potential algorithmic bias and predictors of significant grading divergence through interaction effects. Using records from assessment of professional competence (APC), descriptive and inferential statistical modelling techniques were analyzed. The findings showed significant AI-human grading discrepancy across performance levels and subject areas but no significant gender-based algorithmic grading bias was seen. Likewise, proficiency categories showed no significant influence. AI grades exceed human grades at lower performance levels, while human grades were relatively higher than AI grades at higher levels, implying systemic disagreement at extreme performance level. The strongest predictor of grading discrepancy were performance level and subject area, while subject-specific and calibration effect emerged as the strongest contextual source of grading divergence. Further analysis revealed that higher AI confidence is linked with lower probabilities of grading mismatch, but the predictive strength varied significantly across subjects. Hence, suggests confidence calibration is not evenly reliable across assessment contexts rather it is context dependent. Overall findings imply that AI grading systems may show acceptable overall validity while consistently hiding contextual calibration. The study calls for institutions to build resilient infrastructure and ensure responsible AI adoption in educational practices while seeking a more reliable, scalable, and consistent evaluation to ensure equitable judgment. The findings advance calibration theory, measurement invariance logic, and conditional fairness perspectives.
No comments yet — start the discussion below.