PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 2, 20240 citationsOpen Access

Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?

View Full Paper
NBNishant BalepurRRRachel Rudinger

Key Points

Key points are not available for this paper at this time.

Abstract

Recent work shows that large language models (LLMs) can answer multiple-choice questions using only the choices, but does this mean that MCQA leaderboard rankings of LLMs are largely influenced by abilities in choices-only settings? To answer this, we use a contrast set that probes if LLMs over-rely on choices-only shortcuts in MCQA. While previous works build contrast sets via expensive human annotations or model-generated data which can be biased, we employ graph mining to extract contrast sets from existing MCQA datasets. We use our method on UnifiedQA, a group of six commonsense reasoning datasets with high choices-only accuracy, to build an 820-question contrast set. After validating our contrast set, we test 12 LLMs, finding that these models do not exhibit reliance on choice-only shortcuts when given both the question and choices. Thus, despite the susceptibility~of MCQA to high choices-only accuracy, we argue that LLMs are not obtaining high ranks on MCQA leaderboards just due to their ability to exploit choices-only shortcuts.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Balepur et al. (2024) studied this question.

synapsesocial.com/papers/68e61b7fb6db6435875ae443https://doi.org/10.48550/arxiv.2407.01992
Ask AI
Helpful
Bookmark
Share
View Full Paper