Introduction: Crush injuries are a significant concern in disaster scenarios such as earthquakes. These injuries may lead to serious complications such as crush syndrome and compartment syndrome. Large language models, which enable rapid access to information, are anticipated to serve as valuable information sources for both patients and physicians, particularly in the management of these injuries that are frequently encountered during disaster periods. This study aims to evaluate the performance of four AI-derived LLMs, ChatGPT-3.5, ChatGPT-4, Bing AI, and Google Gemini, in providing accurate and reliable information about crush injuries. Methods: A total of 34 questions related to crush injuries were formulated, 17 of which were designed to simulate inquiries from healthcare professionals and 17 from patients directed toward LLMs. The responses were evaluated by three physiatrists using a 3-point Likert scale as follows: (1) accurate and sufficient, (2) partially accurate but sufficient, and (3) inaccurate and insufficient. In addition, the references provided by the models were analyzed in terms of their accuracy and relevance. Categorical variables were compared using the chi-squared test with Bonferroni correction, and group differences in non-parametric data were analyzed using the Kruskal–Wallis test. Results: When all 34 questions were analyzed together, ChatGPT-3.5 and Google Gemini each provided accurate and sufficient answers to 88.23% (n=30) of the questions, while ChatGPT-4 provided such answers to 82.35% (n=28), and Bing AI to only 55.88% (n=19). Bing AI had the highest mean Likert score among the models, indicating the lowest performance (p=0.002). The number of correct references was higher for ChatGPT-4 and ChatGPT-3.5 than the others (p<0.001), whereas ChatGPT-3.5 and Bing AI provided significantly more hallucinated references compared to Google Gemini (p=0.01 and p=0.03). Conclusion: These findings suggest that LLMs can be valuable tools in managing crush injuries in disaster situations, though caution is needed due to occasional inaccuracies and hallucinated references.
Engin et al. (Tue,) studied this question.