PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 19, 2025Scientific Reports42 citationsOpen Access

”My AI is Lying to Me”: User-reported LLM hallucinations in AI mobile apps reviews

View Full Paper
RMRhodes MassenonIGIshaya GamboJKJaved Ali Khan

Key Points

  • User-reported issues of llm hallucinations account for about 1.75% of reviews flagged for ai errors, impacting user trust.
  • Analysis of 3 million app reviews identified 20,000 candidate reviews, with 1,000 annotated to characterize llm hallucinations.
  • Mixed-methods approach was employed, encompassing a heuristic detection algorithm and linguistic pattern analysis.
  • Findings highlight the importance of understanding llm hallucinations for improving ai mobile app development and user trust.

Abstract

Large Language Models (LLMs) are increasingly integrated into AI-powered mobile applications, offering novel functionalities but also introducing the risk of "hallucinations" generating plausible yet incorrect or nonsensical information. These AI errors can significantly degrade user experience and erode trust. However, there is limited empirical understanding of how users perceive, report, and are impacted by LLM hallucinations in real-world mobile app settings. This paper presents a large-scale empirical study analyzing 3 million user reviews from 90 diverse AI-powered mobile apps to characterize these user-reported issues. Using a mixed-methods approach, a heuristic-based User-Reported LLM Hallucination Detection algorithm were applied to identify 20,000 candidate reviews, from which 1,000 are manually annotated. This analysis estimates the prevalence of user reports indicative of LLM hallucinations, which was found to be approximately 1.75% within reviews initially flagged as relevant to AI errors. A data-driven taxonomy of seven user-perceived LLM hallucination types, were developed with Factual Incorrectness (H1) emerged as the most frequently reported type, accounting for 38% of instances, followed by Nonsensical/Irrelevant Output (H3) at 25%, and Fabricated Information (H2) at 15%. Furthermore, linguistic patterns were identified using N-grams generation, Non-Negative Matrix Factorization (NMF) topics and sentiment characteristics using VADER, showing significantly lower scores for hallucination reports associated with these reviews. These findings offer critical implications for software quality assurance, highlighting the need for targeted monitoring and mitigation strategies for AI mobile apps. This research provides a foundational, user-centric understanding of LLM hallucinations, paving the way for improved AI model development and more trustworthy mobile applications.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Massenon et al. (2025) studied this question.

synapsesocial.com/papers/68af494dad7bf08b1ead4d90https://doi.org/10.1038/s41598-025-15416-8
Ask AI
Helpful
Bookmark
Share
View Full Paper