Randomized trial compares grading accuracy of GPT-4 and human teachers in essays, suggesting AI improvements are needed.
Key Points
GPT-4 shows a risk-averse grading pattern, indicating limitations in adapting to nuanced criteria.
Interrater reliability between GPT-4 and human raters is low, underscoring the challenges in AI assessments.
Assessment of grading methods involved 60 political science essays evaluated by GPT-4 and human educators alike as comparators for accuracy in grading outcomes in higher education context. May enable improved AI applications in educational settings compared to traditional grading methods.