PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 13, 20250 citationsOpen Access

HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems

View Full Paper
SASpandan AnaokarSGShrey GanatraHKHarshvivek Kashid

Key Points

  • HalluDetect achieves an F1 score of 69%, successfully outpacing baseline methods.
  • AgentBot, one of the architectures tested, minimizes hallucinations to 0.4159 per turn, with a token accuracy of 96.13%.
  • The study benchmarks five chatbot architectures, demonstrating effectiveness in mitigating hallucinations.
  • Findings highlight the potential for enhanced factual accuracy in LLM-driven assistants, applicable across various high-risk sectors.

Abstract

Large Language Models (LLMs) are widely used in industry but remain prone to hallucinations, limiting their reliability in critical applications. This work addresses hallucination reduction in consumer grievance chatbots built using LLaMA 3.1 8B Instruct, a compact model frequently used in industry. We develop HalluDetect, an LLM-based hallucination detection system that achieves an F1 score of 69% outperforming baseline detectors by 25.44%. Benchmarking five chatbot architectures, we find that out of them, AgentBot minimizes hallucinations to 0.4159 per turn while maintaining the highest token accuracy (96.13%), making it the most effective mitigation strategy. Our findings provide a scalable framework for hallucination mitigation, demonstrating that optimized inference strategies can significantly improve factual accuracy. While applied to consumer law, our approach generalizes to other high-risk domains, enhancing trust in LLM-driven assistants. We will release the code and dataset

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Anaokar et al. (2025) studied this question.

synapsesocial.com/papers/68ecfebf950606aabec0952dhttps://doi.org/10.48550/arxiv.2509.11619
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation2024 · 3 citations
  2. 2LLM Hallucination: The Curse That Cannot Be Broken2025
  3. 3Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs2024
  4. 4Cost-Effective Hallucination Detection for LLMs2024
  5. 5DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models2024