PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 28, 20250 citationsOpen Access

Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering

View Full Paper
ENEviatar NachshoniACArie CattanSAShmuel Amar

Key Points

  • Models struggle to handle conflicting answers in multi-answer question answering tasks, displaying fragility.
  • Evaluation of eight high-end large language models on a new benchmark shows flawed conflict resolution strategies.
  • The research introduces NATCONFQA, a conflict-aware multi-answer question answering benchmark with detailed labels.
  • Developing conflict-aware datasets through a cost-effective methodology is vital for advancing question answering research.

Abstract

Large Language Models (LLMs) have demonstrated strong performance in question answering (QA) tasks. However, Multi-Answer Question Answering (MAQA), where a question may have several valid answers, remains challenging. Traditional QA settings often assume consistency across evidences, but MAQA can involve conflicting answers. Constructing datasets that reflect such conflicts is costly and labor-intensive, while existing benchmarks often rely on synthetic data, restrict the task to yes/no questions, or apply unverified automated annotation. To advance research in this area, we extend the conflict-aware MAQA setting to require models not only to identify all valid answers, but also to detect specific conflicting answer pairs, if any. To support this task, we introduce a novel cost-effective methodology for leveraging fact-checking datasets to construct NATCONFQA, a new benchmark for realistic, conflict-aware MAQA, enriched with detailed conflict labels, for all answer pairs. We evaluate eight high-end LLMs on NATCONFQA, revealing their fragility in handling various types of conflicts and the flawed strategies they employ to resolve them.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nachshoni et al. (2025) studied this question.

synapsesocial.com/papers/68d913a34ddcf71ba560b7behttps://doi.org/10.48550/arxiv.2508.12355
Ask AI
Helpful
Bookmark
Share
View Full Paper