Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment
Quick Answer
This study introduces a generative AI-assisted summarization framework using GPT-5 variants to enhance automated essay scoring (AES) while addressing transformer input-length limitations.
Quick Take
The GPT-5 mini variant achieved the highest agreement with human ratings, indicating a promising approach for scalable writing assessment, although higher-scoring essays showed reduced summary quality due to complexity.
Key Points
- GPT-5 mini achieved the highest agreement with human ratings in AES evaluations.
- Summarization quality decreases for higher-scoring essays, indicating complexity challenges.
- The framework integrates handcrafted linguistic features with summary representations.
- Evaluation metrics include quadratic weighted kappa and various summary quality metrics.
- Generative AI summarization shows promise for scalable educational assessments.
DeepSignal Analysis
What happened
A study proposed a generative AI-assisted summarization framework using GPT-5 variants to enhance automated essay scoring (AES). The framework addresses transformer input-length limitations, which can lead to information loss in long essays. The GPT-5 mini variant showed the highest agreement with human ratings, while the quality of summaries decreased for more complex, higher-scoring essays.
Key evidence
- The study utilized the ASAP 2.0 dataset to evaluate the performance of three GPT-5 variants: GPT-5, GPT-5 mini, and GPT-5 nano.
- Scoring reliability was measured using quadratic weighted kappa (QWK), while summary quality was assessed through various metrics including lexical overlap and semantic similarity.
- Results indicated that the GPT-5 mini variant achieved the highest agreement with human ratings, suggesting its effectiveness for AES.
Why it matters
This research highlights the potential of generative AI in educational assessment, particularly in improving the reliability of automated essay scoring. By addressing input-length limitations, the proposed framework could facilitate more accurate evaluations of student writing. However, the observed trade-offs between summary quality and essay complexity raise concerns about the preservation of critical writing signals, which are essential for fair assessment.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Automated essay scoring (AES) enables scalable assessment and timely feedback but remains challenged by transformer input-length limitations, which can cause information loss when processing long essays. This study proposes a generative AI-assisted summarization framework to improve long-form essay representation while maintaining scoring reliability. Using the ASAP 2.0 dataset, we generate controlled-length summaries with three GPT-5 variants (GPT-5, GPT-5 mini, and GPT-5 nano) and use them as inputs for downstream AES models. To preserve original writing signals, handcrafted linguistic features extracted from full essays are integrated with summary representations to form a hybrid framework. The approach is evaluated in terms of scoring performance, summarization quality, and computational cost. Scoring reliability is measured using quadratic weighted kappa (QWK), while summary quality is assessed through lexical overlap, semantic similarity, information retention, and redundancy metrics. Results show that GPT-5 mini achieves the highest agreement with human ratings, whereas GPT-5 produces the strongest summarization quality. Summary quality decreases for higher-scoring essays, indicating that more complex writing is more difficult to compress without information loss. These findings reveal trade-offs among model capacity, summary fidelity, cost efficiency, and preservation of educational constructs. This study provides an initial controlled evaluation of GPT-based summarization for AES and identifies important baselines and ablation studies required for future generalization. Overall, generative AI summarization offers a promising approach for scalable writing assessment while requiring careful validation of information preservation and fairness.
| Comments: | 23 pages, 7 figures, 5 tables |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.15829 [cs.CL] |
| (or arXiv:2607.15829v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15829 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haowei Hua [view email]
[v1]
Fri, 17 Jul 2026 10:39:58 UTC (4,038 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.