Cheng Tang, Jing Li, Jiawei Xiong, George Engelhard · The Journal of Experimental Education 2026 · 2026
DOI: 10.1080/00220973.2026.2730125
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large Language Models (LLMs) have shown promise in automated writing evaluation that provides estimates of writing proficiency ratings and feedback, yet validation studies concentrate on score agreement, overlooking whether human and LLM raters attend to the same aspects in rubrics. We evaluate comparability between a human reference benchmark and an LLM rater using both agreement indices and Many-Facet Rasch Measurement (MFRM) in ratings and feedback, and propose the psychometric approach that separates what theme the rater emphasizes in feedback and how strongly the rater evaluates within a theme. We use an argumentative writing assessment with 300 examinees, scored by human raters and an LLM under multiple prompting conditions. For feedback, 26 rubric codes are mapped into five theory-grounded themes and modeled via dichotomous MFRM for theme detection and partial-credit MFRM for code-level severity. Results show that the surface agreement was improved with few-shot prompting, but systematic rater differences persist in where score and feedback attention are allocated and how severity is applied across themes based on a stable rater-by-domain interaction effect, indicating structured divergence. These findings demonstrate why agreement is insufficient and illustrate how a score and feedback-focused MFRM can identify systematic discrepancies that matter for assessment result interpretation.
No comments yet — start the discussion below.