Authors
Loading...
Benchmark assesses gender classification and vote prediction in large language models, highlighting bias.
Yang et al. (2025) studied this question.