Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
January 1, 2024Open Access

Humans or LLMs as the Judge? A Study on Judgement Bias

View Full Paper
Ask AI
Bookmark
Share

Authors

GCGuiming Hardy ChenSCShunian ChenZLZiche Liu

Discussion

Loading...

Member takes

Overview

Empirical analysis reveals substantial judgment biases in both human and large language model evaluators, highlighting the need for robust automated assessment systems.

Key Points

  • Quantify and compare cognitive and systematic judgment biases—including misinformation oversight, gender, authority, and beauty biases—in human and large language model evaluators without relying on ground-truth annotations.
  • Curated an evaluation benchmark grounded in the revised Bloom's Taxonomy.
  • Conducted thousands of evaluations testing human and large language model judges against systematic perturbations across four bias dimensions.
  • Designed adversarial attack protocols exploiting identified evaluation biases to deliberately manipulate large language model scoring outcomes.
  • Both human and cutting-edge large language model judges exhibited measurable vulnerability to perturbations across misinformation oversight, gender, authority, and beauty biases.
  • Identified vulnerabilities enabled successful adversarial attacks that systematically skewed evaluation judgments produced by large language models.

Cite This Study

Chen et al. (2024) studied this question.

synapsesocial.com/papers/69d6bceff174babf6cab3550https://doi.org/10.18653/v1/2024.emnlp-main.474
View Full Paper
Ask AI
Bookmark
Share