Aim: Evaluataion of the diagnostic accuracy and clinical utility of the Medicine GPT model in identifying and managing rare diseases in the emergency department. Materials and methods: A retrospective study was conducted using 100 published rare case reports retrieved from PubMed. Cases were selected based on three common complaints: chest pain (n = 50), abdominal pain (n = 30), and shortness of breath (n = 20). Each case’s clinical presentation was input into Medicine GPT, which then generated a leading diagnosis. The primary outcome was the proportion of correct diagnoses. Cohen’s Kappa (κ) was calculated to determine inter-rater reliability. Diagnostic accuracy varied by symptom category: chest pain (92%, κ = 0.84), abdominal pain (87%, κ = 0.81), and shortness of breath (90%, κ = 0.875). Based on full clinical presentations, overall diagnostic agreement reached 90% (κ = 0.84), indicating substantial to almost perfect consistency. Conclusions: Medicine GPT demonstrated substantial diagnostic accuracy for rare emergency department presentations, indicating potential to enhance clinical decision-making. It may serve as a foundation for refining future AI models in real-world settings, but always should remain only as an assistive tool for professional clinical judgment.
Obeid et al. (Mon,) studied this question.