PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 22, 20243 citations

A Comparative Analysis of Different Large Language Models in Evaluating Student-Generated Questions

View Full Paper
ZMZejia MiKLKangkang Li

Key Points

Key points are not available for this paper at this time.

Abstract

Student-generated questions (SGQs) have proven to be a meaningful learning tool, fostering advanced thinking skills in students and aiding teachers in understanding student learning progress. However, grading the quality of SGQs demands significant effort from teachers. In this study, we explore the suitability of large language models in evaluating SGQs and identify which models can effectively replace expert evaluation of practical teaching problems. We devised a five-dimension scale, using expert ratings as the gold standard, and employed Kendall's W consistency analysis to systematically compare different large language model evaluations against expert ratings from six aspects of the scale. The research confirmed the applicability of large language models (LLMs) for the evaluation of SGQs and the exceptional performance of ChatGPT 4.0, which can assist experts in evaluating SGQs. This study aims to facilitate the implementation of artificial intelligence generated content (AIGC) in education and reinforces the belief in the substantial potential of large language models for future applications and research in the field of education.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mi et al. (2024) studied this question.

synapsesocial.com/papers/68e72cd4b6db6435876a6258https://doi.org/10.1109/iceit61397.2024.10540914
Ask AI
Helpful
Bookmark
Share
View Full Paper