Empirical analysis reveals substantial judgment biases in both human and large language model evaluators, highlighting the need for robust automated assessment systems.
Key Points
Quantify and compare cognitive and systematic judgment biases—including misinformation oversight, gender, authority, and beauty biases—in human and large language model evaluators without relying on ground-truth annotations.
Curated an evaluation benchmark grounded in the revised Bloom's Taxonomy.
Conducted thousands of evaluations testing human and large language model judges against systematic perturbations across four bias dimensions.
Designed adversarial attack protocols exploiting identified evaluation biases to deliberately manipulate large language model scoring outcomes.
Both human and cutting-edge large language model judges exhibited measurable vulnerability to perturbations across misinformation oversight, gender, authority, and beauty biases.
Identified vulnerabilities enabled successful adversarial attacks that systematically skewed evaluation judgments produced by large language models.