Fatima Ismail · · 2026
DOI: 10.35542/osf.io/7c62b_v2
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Automated evaluation of IELTS Writing Task 2 essays requires judging task fulfilment, coherence, vocabulary, and grammar in ways that approximate trained human examiners, while remaining fast and inexpensive enough for practical deployment. This paper presents a systematic comparison of six architectures for automated IELTS essay evaluation a local generative LLM baseline (Gemma 2 9B), the same model augmented with rubric-grounded and example-grounded Retrieval-Augmented Generation (RAG), a multi-stage orchestration pipeline, a cloud-hosted LLM (Llama 3.3 70B), and a fine-tuned discriminative transformer (DeBERTa-v3-base) applied to a common 296-word benchmark essay and the same four official IELTS criteria. Rather than treating the highest observed score as evidence of accuracy, we compare scoring behavior, consistency, latency, token usage, and estimated commercial cost. Observed overall scores ranged from 6.3 to 8.5 on the 9-band scale; rubric grounding shifted the ungrounded baseline's score substantially downward (8.5 → 6.3); multi-stage orchestration improved cross-stage consistency but incurred roughly 23 minutes of total latency; and the cloud LLM offered the best latency-cost trade-off among generative approaches. The fine-tuned DeBERT a model combined the highest observed score in the final comparison with the fastest local inference (1.33 seconds) and no per-request cost. We argue that scoring reliability and explanatory quality are distinct capabilities that need not come from the same model, and propose a hybrid architecture pairing a specialized scorer with a generative LLM for feedback generation. Keywords: automated essay scoring; IELTS; large language models; retrieval-augmented generation; finetuned transformers; multi-agent evaluation; educational AI
No comments yet — start the discussion below.