PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 8, 20250 citationsOpen Access

Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems

View Full Paper
AAAakriti AgrawalRARohith AralikattiASAnirudh Satheesh

Key Points

  • The proposed approach improves reasoning by leveraging log-likelihood scores from multiple LLMs, and knowledge of their outputs.
  • Improvements of 4%, 3%, and 5% were demonstrated in both debate and non-debate settings, with key datasets being GSM8K, MMLU, and ARC.
  • Assessment involved assessing diverse responses in multi-LLM interactions rather than single LLM models, focusing on efficiency and accuracy.
  • Enhanced response selection may enable broader applications for complex decision-making tasks with LLMs.

Abstract

Large Language Models (LLMs) have demonstrated exceptional capabilities, yet selecting the most reliable response from multiple LLMs remains a challenge, particularly in resource-constrained settings. Existing approaches often depend on costly external verifiers, human evaluators, or self-consistency techniques that require multiple samples from a single model. While multi-LLM systems produce more diverse responses than single models and thus have greater potential, they often underperform compared to single LLM self-consistency. We propose a principled, novel and computationally efficient method to select the best response from multiple different LLMs using a calibrated log-likelihood score, implicitly leveraging the inherent knowledge and confidence of these models. Our method demonstrates improvements of approx. 4%, 3%, and 5% across both debate (multi-round LLM discussions) and non-debate (Best-of-N with multiple LLMs) settings on GSM8K, MMLU (6 subsets), and ARC datasets respectively.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Agrawal et al. (2025) studied this question.

synapsesocial.com/papers/690e8b75a5b062d7a4e73883https://doi.org/10.48550/arxiv.2510.02377
Ask AI
Helpful
Bookmark
Share
View Full Paper