Abstract—Large Language Models (LLMs) offer a scalable way to analyse classroom practice, but they also risk amplifying existing gender biases. This paper presents a fairness-aware audit and calibration pipeline for an LLM-based system that classifies teacher questions according to the Singapore Teaching Practice (STP). We analyse classroom audio from eight secondary-school teachers (4 male, 4 female) using Whisper-medium (WER ≈ 0.04) and GPT-4o-mini to label “deepening” versus other questions. A baseline audit reveals an F1 score of 0.79 for male teachers and 0.71 for female teachers (gap = 0.08), driven mainly by lower recall for female voices. We then apply post-hoc calibration, combining Platt scaling with group-specific decision thresholds inspired by Equalized Odds, reducing the F1 gap to 0.01 without retraining the model. We argue that such fairnessaware calibration is a practical prerequisite for the ethical use of generative AI in teacher evaluation. Index Terms—Large Language Models (LLMs), Teacher Assessment, Post-hoc Calibration, Gender Bias
Abidi et al. (Mon,) studied this question.