Facial recognition systems are increasingly deployed in law enforcement and security contexts, where algorithmic decisions can carry significant societal consequences. Despite high reported accuracy, growing evidence demonstrates that such systems often exhibit uneven performance across demographic groups, leading to disproportionate error rates and potential harm. This paper argues that aggregate accuracy is an insufficient metric for evaluating the fairness and reliability of facial recognition systems for high-stakes environments. Through analysis of subgroup-level error distribution, including false positive and false negative rates, we demonstrate how overall performance metrics can obscure critical disparities across demographic groups. Drawing on existing literature and empirical observations from classification-based systems, the paper highlights the operational risks associated with accuracy-centric evaluation practices, particularly in law enforcement applications where misclassification may result in wrongful suspicion or missed identification. We further discuss the importance of model-agnostic fairness auditing approaches that enable post-deployment evaluation without access to proprietary systems. Finally, the paper outlines the inherent trade-offs between fairness and accuracy and emphasizes the need for more comprehensive fairness-aware evaluation strategies in high-stakes AI systems.
Khalid Adnan Alsayed (2026) studied this question.