Tahereh Firoozi, Doris Abroampah, Hamid Mohammadi, Mark Gierl · International Journal of Testing 2026 · 2026
DOI: 10.1080/15305058.2026.2730013
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Automatic essay scoring is an educational testing technology where written-response tasks are scored using a computer program. One significant challenge that must be addressed when implementing automatic essay scoring in a multilingual context resides with the issue of data scarcity. Large samples of multilingual essays are typically not available to train AES systems. The purpose of our study was to address the data scarcity challenge found commonly in multilingual settings by describing and implementing three data augmentation methods—paraphrasing, synonym replacement, and random deletion—using of four commonly used large language models—mBERT, RoBERTa, LaBSE, and DistilmBERT—in order to evaluate their performance with essays written in German, Italian, and Czech. The results demonstrated that paraphrasing using the LaBSE large language model consistently produced the most accurate and consistent results across the three language groups. Directions for future research using different scoring rubrics and different sample sizes are also discussed.
No comments yet — start the discussion below.