Haowei Hua · Next research. 2026 · 2026
DOI: 10.1016/j.nexres.2026.102442
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Automated essay scoring (AES) has become an important tool in educational assessment by enabling scalable evaluation of student writing and timely formative feedback. However, many transformer-based embedding models (e.g., BERT, RoBERTa, DeBERTa) are constrained by strict input-length limits, making it difficult to process long-form essays without truncation and potential loss of educationally relevant information. This study proposes a generative AI-assisted summarization framework to address this limitation while preserving scoring reliability. Using the ASAP 2.0 dataset, we generate controlled-length summaries of student essays with three GPT-5 model variants (GPT-5, GPT-5 mini, GPT-5 nano) and use these summaries as inputs for downstream scoring models. To retain important linguistic signals, we integrate handcrafted features from the original essays, forming a hybrid representation of writing. The framework is evaluated across downstream scoring performance, summarization quality, and model-cost considerations. Scoring performance is measured using quadratic weighted kappa (QWK), while summarization quality is assessed through lexical overlap, semantic similarity, information retention, and redundancy. Results show that GPT-5 mini achieves the highest agreement with human raters, while GPT-5 provides the strongest summarization quality. Summarization performance declines for higher-scoring essays, suggesting that more complex writing is harder to compress without information loss. These findings highlight trade-offs among model capacity, summary fidelity, computational cost, and construct preservation. The study should be interpreted as an initial controlled evaluation of GPT-based summarization for AES rather than a complete benchmark against all long-context and direct-scoring alternatives. Overall, the study demonstrates that generative AI summarization can support reliable and scalable assessment of writing ability in educational contexts, while also identifying the baseline and ablation experiments needed for stronger generalization claims.
No comments yet — start the discussion below.