Abstract—Automated marking has emerged as a promising solution to enhance efficiency, scalability, and feedback quality in education. While multiple-choice questions have long benefited from automated grading, the evaluation of long-form answers has remained challenging until recent advances in Natural Language Processing and Generative AI. This paper proposes a practical architecture for automated marking using locally hosted large language models (LLMs) and open-source tools, specifically Ollama and AnythingLLM, deployed on a standard laptop. A dataset of 149 submissions marked by three human assessors was used to train and evaluate the system. Results show that the AI agent, powered by Qwen3, produced average scores within acceptable moderation thresholds and generated more detailed feedback, offering richer formative insights. Despite limitations in response time and scalability, the study demonstrates the feasibility of running LLMs locally for automated marking and highlights potential applications beyond grading, including moderating discrepancies between markers, generating instant feedback, and teaching prompt engineering. Future work will explore fine-tuned smaller models and the extension of automated marking to more complex assessments such as long-form reports. Index Terms—Automated marking, generative AI, LLM
Lim et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: