PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 2, 2026BMJ Quality & Safety0 citationsOpen Access

AI-driven analysis of patient safety reports using large language models: an exploratory multiple methods study

View Full Paper
KCKevin ChenKRKiley RogersWHWilliam J. Haberkorn

Key Points

  • This study aims to evaluate an AI-driven approach using large language models to identify patient safety issues and trends in reports.
  • Evaluated OpenAI’s GPT-4 model accuracy for analyzing safety event reports.
  • Extracted data from 9357 free-text narratives to create a taxonomy of safety issues.
  • Conducted validation reviews with patient safety experts on report subsets.
  • Assessed stakeholder feedback through 10 interviews regarding the dashboards presented.
  • LLM achieved 94% mean agreement among reviewers on identifying problems in reports.
  • Agreement rates of 91.5% for 'parent' and 83.3% for 'child' category labels were recorded.
  • Identified previously hidden patterns of patient safety issues.
  • Stakeholders found insights clear and valuable, perceiving minimal barriers to adoption.

Abstract

Introduction Patient safety event reporting systems are widely used, yet organisations face challenges analysing the high volume of incident reports. While low-harm events represent the majority of submissions, they are rarely examined systematically due to the time and resources required to manually review the complex, lengthy, narrative data. Emerging technology like large language models (LLMs) offers new opportunities to address this gap. This study aimed to develop and evaluate an artificial intelligence-driven approach to identify patient safety issues, uncover system-level trends and assess readiness for implementation within a US healthcare system. Methods We quantitatively evaluated OpenAI’s GPT-4o model accuracy in analysing patient safety event reports and qualitatively assessed pre-implementation outcomes. The LLM extracted safety problems from 9357 free-text narratives and then generated a taxonomy of ‘parent’ and ‘child’ subcategories. The model labelled every report’s problem list using the taxonomy, and dashboards were developed to visualise trends. Patient safety experts reviewed two separate subsets of reports (n=100 and n=219) to validate the LLM’s accuracy. We conducted 10 stakeholder interviews to assess the dashboards’ acceptability, appropriateness and adoption. Results The LLM had scores of 94% mean agreement among reviewers in identifying problems, and 91.5% and 83.3% agreement in assigning ‘parent’ and ‘child’ category labels, respectively. The model identified previously hidden patterns of patient safety issues. Stakeholders described LLM-generated insights as clear, appropriate and valuable for their work. They perceived few barriers to adoption and believed the model could expedite manual report reviews and support system-level quality improvement. Conclusion LLMs offer an approach to capture and analyse multidimensional concepts in patient safety reports. By extracting problem summaries and categorising them into an accepted taxonomy, they can expose previously unidentified trends and provide visibility into system-level risks. LLMs could effectively augment and expedite manual reviews of safety events, while guiding prioritisation of quality and safety interventions.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chen et al. (2026) studied this question.

synapsesocial.com/papers/6980fe48c1c9540dea810273https://doi.org/10.1136/bmjqs-2025-019495
Ask AI
Helpful
Bookmark
Share
View Full Paper