Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 12, 2025Open Access

Benchmarking Gender and Political Bias in Large Language Models

View Full Paper
Ask AI
Bookmark
Share

Authors

JYJinrui YangXHXudong HanTBTimothy Baldwin

Discussion

Loading...

Member takes

Overview

Benchmark assesses gender classification and vote prediction in large language models, highlighting bias.

Key Points

  • Large language models frequently misclassify female Members of the European Parliament as male, indicating a systematic bias.
  • Evaluation shows that LLMs tend to favor centrist political groups while exhibiting reduced accuracy on far-left and far-right categories.
  • Proprietary models like GPT-4o outperform open-weight alternatives in robustness, fairness, and accuracy on politically sensitive tasks.
  • The EuroParlVote dataset provides essential data for future research on fairness and accountability in natural language processing.

Cite This Study

Yang et al. (2025) studied this question.

synapsesocial.com/papers/68ebffcfdef9fcb308ff2666https://doi.org/10.48550/arxiv.2509.06164
View Full Paper
Ask AI
Bookmark
Share