PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 29, 2026British journal of surgery0 citations

SRS54 - How have generic large language models progressed in their ability to write clinical letters and manage patients in the virtual fracture clinic?

View Full Paper
ASA SmithJBJames B. BrockRARana Anss

Key Points

  • The study evaluates how large language models evolved in creating clinical letters and management plans for orthopaedic cases.
  • Generated fifteen clinical scenarios for evaluation.
  • Prompts given to GPT-4, ChatGPT, and GPT-3 to write clinic letters and management plans.
  • Assessed readability using the Readable Tool and accuracy by six blinded orthopaedic consultants.
  • GPT-4 showed the highest readability scores among the models tested.
  • All models generated accurate letters with GPT-4 and Chat-GPT outperforming GPT-3 significantly.
  • The accuracy of management plans increased from GPT-3 to GPT-4, demonstrating statistical significance.

Abstract

Abstract Objective to evaluate the progression of large language models (LLMs) and their ability to write clinic letters and management plans for common orthopaedic scenarios. Methods Fifteen clinical scenarios were generated and GPT-4, Chat-GPT and GPT-3 were single prompted to write clinic letters and management plans. Letters were assessed for readability using the Readable Tool. Accuracy of letters and management plans were assessed by six independent blinded orthopaedic consultants. Results Readability was compared using Flesch-Kincade Grade Level (GPT-4:9.11;(SD 0.98);ChatGPT:8.77 (SD 0.918);GPT-3:8.47 (SD 0.982)), Flesch Readability Ease (GPT-4:34.26 (SD 7.91);ChatGPT:58.2 (SD 4.00);GPT-3,59.3 (SD 6.98)). GPT-4, Chat-GPT and GPT-3 produced accurate letters (Mean = 8.75/10 (SD 0.96), 8.7/10 (SD 0.60), 7.3/10 (SD 1.41)) respectively. GPT4 and Chat-GPT had a significantly increased letter accuracy compared to GPT-3 (P = 0.024, P = 0.019). Consultant-rated accuracy comparisons across 4.0, 3.5 and 3.0 revealed that ChatGPT-4 exhibited the highest accuracy for management plans (9.08/10 95%c.i., 8.25–9.9). This represents a statistically significant progression of the ability of a large language model to provide accurate management plans from GPT-3 6.84 (95% c.i., 5.41–8.27), to ChatGPT 7.63 to GPT4 (P 0.0001). Conclusions This study shows that next generation LLMs are effective for generation of clinic letters which are readable and accurate. Further, LLMs can produce generic management plans that are often accurate, demonstrating their evolving improvement. Given these findings a specific LLM trained on accurate and secure healthcare data could be an excellent streamlining tool for clinicians in high demand areas such as virtual fracture clinics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Smith et al. (2026) studied this question.

synapsesocial.com/papers/69c8c2e4de0f0f753b39d6c5https://doi.org/10.1093/bjs/znag018.066
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Application of generative language models to orthopaedic practice2024 · 27 citations
  2. 2Systematic Review on Large Language Models in Orthopaedic Surgery2025 · 4 citations
  3. 3Evaluating the progression of artificial intelligence and large language models in medicine through comparative analysis of ChatGPT-3.5 and ChatGPT-4 in generating vascular surgery recommendations2023 · 12 citations
  4. 4The Temporal Evolution of Large Language Model Performance: A Comparative Analysis of Past and Current Outputs in Scientific and Medical Research2025
  5. 5ENHANCING INFORMED CONSENT IN ORTHOPAEDIC SURGERY: A PROOF-OF-CONCEPT STUDY OF LANGUAGE MODEL-GENERATED CLINIC LETTERS2025