Key points are not available for this paper at this time.
Machine learning-based automated essay scoring (AES) and feedback generation (AFG) tools have been developed since the 1960s, with some commercially deployed. Such systems are expected to lighten the labor-intensive task of essay scoring and help provide timely feedback to students. However, these tools have not been used effectively enough in practice despite the long history of this research field. We aim to determine how accurately currently available models can make the necessary corrections and detect improvable segments of text. Latest attempts of AFG utilize state-of-the-art generative large language models (LLMs) for writing evaluation. To the best of our knowledge, however, none of the past studies have included generative LLM-based models in a fine-grained sentence-to-sentence comparison with human feedback or among AI tools. To fill these gaps, we conduct an experimental comparison of human feedback and three AI tools developed at different stages of technological advancement. Findings indicate that the overlap between human and AI feedback is predominantly limited to surface-level linguistic features and that generative AI-augmented tools demonstrate a markedly higher capability than a tool based on conventional rule-based AI.
Sasaki et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: