PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 7, 20243 citationsOpen Access

HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild

View Full Paper
ZZZhiying ZhuZSZhiqing SunYYYiming Yang

Key Points

Key points are not available for this paper at this time.

Abstract

Hallucinations pose a significant challenge to the reliability of large language models (LLMs) in critical domains. Recent benchmarks designed to assess LLM hallucinations within conventional NLP tasks, such as knowledge-intensive question answering (QA) and summarization, are insufficient for capturing the complexities of user-LLM interactions in dynamic, real-world settings. To address this gap, we introduce HaluEval-Wild, the first benchmark specifically designed to evaluate LLM hallucinations in the wild. We meticulously collect challenging (adversarially filtered by Alpaca) user queries from existing real-world user-LLM interaction datasets, including ShareGPT, to evaluate the hallucination rates of various LLMs. Upon analyzing the collected queries, we categorize them into five distinct types, which enables a fine-grained analysis of the types of hallucinations LLMs exhibit, and synthesize the reference answers with the powerful GPT-4 model and retrieval-augmented generation (RAG). Our benchmark offers a novel approach towards enhancing our comprehension and improvement of LLM reliability in scenarios reflective of real-world interactions.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhu et al. (2024) studied this question.

synapsesocial.com/papers/68e7555db6db6435876cd1f9https://doi.org/10.48550/arxiv.2403.04307
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models2024
  2. 2HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation2024 · 4 citations
  3. 3Hallucination Detection, Categorization, and Mitigation in Large Language Models: A Cross-Domain Evaluation Framework2026
  4. 4DefAn: Definitive Answer Dataset for LLMs Hallucination Evaluation2024 · 2 citations
  5. 5Deception-Based Benchmarking: Measuring LLM Susceptibility to Induced Hallucination in Reasoning Tasks Using Misleading Prompts2024 · 2 citations