PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 10, 2025Pharmacy3 citationsOpen Access

Accuracy and Safety of ChatGPT-3.5 in Assessing Over-the-Counter Medication Use During Pregnancy: A Descriptive Comparative Study

View Full Paper
BCBernadette CornelisonDADavid R. AxonBAB. Abbott

Key Points

  • Responses generated by ChatGPT-3.5 scored high on correctness, with a median score of 5 out of 5.
  • In terms of completeness, ChatGPT-3.5 achieved a median score of 4, indicating good but not perfect information delivery.
  • Despite high accuracy, safety errors were present in 9% of evaluations, reflecting risk in using AI as a sole resource during pregnancy.
  • The independent ratings showed consensus on ChatGPT-3.5's limitations, emphasizing the need for expert consultation when assessing OTC medications.

Abstract

As artificial intelligence (AI) becomes increasingly utilized to perform tasks requiring human intelligence, patients who are pregnant may turn to AI for advice on over-the-counter (OTC) medications. However, medications used in pregnancy may pose profound safety concerns limited by data availability. This study focuses on a chatbot's ability to accurately provide information regarding OTC medications as it relates to patients that are pregnant. A prospective, descriptive design was used to compare the responses generated by the Chat Generative Pre-Trained Transformer 3.5 (ChatGPT-3.5) to the information provided by UpToDate®. Eighty-seven of the top pharmacist-recommended OTC drugs in the United States (U.S.) as identified by Pharmacy Times were assessed for safe use in pregnancy using ChatGPT-3.5. A piloted, standard prompt was input into ChatGPT-3.5, and the responses were recorded. Two groups independently rated the responses compared to UpToDate on their correctness, completeness, and safety using a 5-point Likert scale. After independent evaluations, the groups discussed the findings to reach a consensus, with a third independent investigator giving final ratings. For correctness, the median score was 5 (interquartile range IQR: 5-5). For completeness, the median score was 4 (IQR: 4-5). For safety, the median score was 5 (IQR: 5-5). Despite high overall scores, the safety errors in 9% of the evaluations (n = 8), including omissions that pose a risk of serious complications, currently renders the chatbot an unsafe standalone resource for this purpose.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cornelison et al. (2025) studied this question.

synapsesocial.com/papers/68c1a12d54b1d3bfb60dc3edhttps://doi.org/10.3390/pharmacy13040104
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Assessment Study of ChatGPT-3.5’s Performance on the Final Polish Medical Examination: Accuracy in Answering 980 Questions2024 · 18 citations
  2. 2Future of ADHD Care: Evaluating the Efficacy of ChatGPT in Therapy Enhancement2024 · 50 citations
  3. 3Artificial Intelligence in Postoperative Care: Assessing Large Language Models for Patient Recommendations in Plastic Surgery2024 · 31 citations
  4. 4Medication Usage Record-Based Predictive Modeling of Neurodevelopmental Abnormality in Infants under One Year: A Prospective Birth Cohort Study2024 · 3 citations
  5. 5Accuracy and Completeness of ChatGPT-Generated Information on Interceptive Orthodontics: A Multicenter Collaborative Study2024 · 93 citations