Cătălin Anghel, Andreea Alexandra Anghel, Marian Crăciun, Antonio Stefan Balau, Adrian Istrate, Adina Cocu, Constantin Adrian Andrei, Aurelian-Dumitrache Anghele · Computers 2026 · 2026
DOI: 10.3390/computers15090571
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large language models are increasingly used for automated grading, but final-score agreement does not reveal whether credit is assigned to the rubric element affected by a response change. This study introduces CreditTrace-LLM, a controlled framework for auditing directional responsiveness, credit locality, paraphrase stability, and correspondence with expert evaluations in rubric-guided grading. The evaluation used 500 response families across ten questions, equally divided between technical and argumentative tasks. Each family included a baseline response (C0), a positive intervention adding rubric-relevant evidence (C+), a negative intervention removing or weakening such evidence (C−), and a placebo paraphrase (CP) constructed through lexical or syntactic reformulation with the intention of preserving rubric-relevant meaning. Eight local open-weight LLM graders produced 15,999 structurally valid Gold-point-level outputs from 16,000 expected evaluations. Targeted Gold-point scores changed in the expected direction in 75.7% of C+ comparisons and 50.2% of C− comparisons. Localized directional success was lower under C+ and C−, at 31.5% and 21.3%, respectively, indicating frequent non-target score changes. Under CP, total-score stability was 75.4%, while complete-profile stability was 73.6%. Direction concordance with the mean expert score change was 69.8% for C+, 47.4% for C−, and 18.2% for CP. The CP condition also showed substantial variability in expert scoring, particularly for argumentative responses, indicating that intended rubric-relevant meaning preservation did not guarantee score invariance. Final-score behavior alone is insufficient for validating rubric-guided LLM grading. CreditTrace-LLM therefore evaluates target responsiveness, credit locality, paraphrase stability, and traceable Gold-point-level outputs as complementary diagnostic dimensions rather than as predefined criteria for acceptable grading performance.
No comments yet — start the discussion below.