PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 28, 2026Scientific Reports0 citationsOpen Access

Evaluating large language model`s performance in answering principles of health course questions

MKMohsen KhosraviEDEmine Kübra DindarBSBurak Sayar

Key Points

  • The aim is to evaluate the performance of multiple large language models in responding to health-related questions.
  • Cross-sectional study conducted in 2025 utilizing ChatGPT-4o, Gemini 2.5, Copilot 2025, and Perplexity 2.250619.0.
  • A confusion matrix was constructed to compare LLM performance, calculating sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and overall accuracy.
  • All LLMs demonstrated perfect sensitivity with a value of 1.
  • ChatGPT and Perplexity achieved the highest accuracy rates of 0.93, while Gemini and Copilot both had an accuracy of 0.86.
  • Performance generally declined with increased complexity and length of questionnaire items, with Copilot showing difficulty with quantitative questions.

Abstract

Introduction The literature highlights the considerable potential of Artificial Intelligence (AI), particularly large language models (LLMs), in advancing health promotion among individuals. This study aimed to evaluate the performance of several LLMs in responding to questions from the Principles of Health course. This cross-sectional study was conducted in 2025. The LLMs evaluated included ChatGPT-4o, Gemini 2.5, Copilot 2025, and Perplexity 2.250619.0. These LLMs were utilized to respond to the study questionnaire pertaining to the Principles of Health course. To analyze and compare the performance of the LLMs in answering the research questions, a confusion matrix was constructed. Accordingly, four key metrics were calculated: sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV), in addition to overall accuracy. The LLMs included in the study demonstrated perfect sensitivity, each achieving a value of 1. Regarding specificity, ChatGPT and Perplexity attained the highest scores of 0.8, while Gemini and Copilot exhibited comparatively lower specificity values of 0.66 and 0.6, respectively. Furthermore, ChatGPT and Perplexity recorded the highest accuracy rates of 0.93, surpassing Gemini and Copilot, both of which achieved an accuracy of 0.86. The findings provided a detailed assessment of the performance of the LLMs. Results indicated that the performance of LLMs generally declined as the complexity, length, and verbosity of questionnaire items increased. Additionally, certain LLMs, such as Copilot, demonstrated particular difficulty when responding to quantitative questions involving numerical data. Further research is recommended to investigate these observations more comprehensively.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Khosravi et al. (2026) studied this question.

synapsesocial.com/papers/69f04e9b727298f751e72869https://doi.org/10.1038/s41598-026-49466-3
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of Large Language Models in Answering Healthcare Delivery Questions: A Quantitative Cross‐Sectional Study2026 · 1 citations
  2. 2Rise of the Machines: Comparing Performance of Artificial Intelligence Large Language Models on Pharmacy Specialty Certification Examination Practice Questions2026 · 1 citations
  3. 3Comparison of Performance of Leading Large Language Models in Answering Medical Pathology Questions in Dentistry Specialization Education Entrance Exams: A Cross-Sectional Research2026
  4. 4Clinical performance and readability evaluation of large language models for patient communication in heart failure and cardiomyopathies2026
  5. 5Clinical Assessment of Large Language Models: A Comprehensive Multi-domain Performance Study for Healthcare Applications2025 · 1 citations