Purpose This study aims to discuss transparency and accountability in hate speech detection and adjacent fields, with a focus on misogyny, by analyzing annotator documentation; as well as to re-examine the prevailing overreliance on performance metrics alone for model evaluation and to emphasize the critical role of annotator documentation in ensuring equitable and trustworthy detection systems. Design/methodology/approach The discussion is based on the analysis of annotator documentation across 25 studies and a post hoc framework of six disclosure dimensions: annotator selection, quantity, training, demographics, compensation and expertise. Building on this framework, this paper introduces a novel weighted annotator metadata transparency (AMT) score that prioritizes demographics and domain expertise as critical factors for equitable model development. Findings The results of this study reveals significant transparency gaps: surveyed articles achieved only a 27% average AMT score, with particularly concerning deficiencies in demographics (20%) and domain expertise (0%). Furthermore, this paper identifies an inverse relationship between transparency and performance: models with higher F1 scores consistently demonstrated poorer annotator documentation, while more transparent studies clustered in the lower performance range. Originality/value This study introduces a novel, weighted framework (the AMT score) for evaluating AMT. The finding of an inverse relationship between reported model performance and transparency is a critical and original contribution that challenges current evaluation paradigms. The authors conclude by providing a set of best practices to guide the research community toward more equitable and trustworthy model development.
Ribeiro et al. (Mon,) studied this question.