PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 24, 20251 citationsOpen Access

Large Language Models for Zero-Shot Procedure Extraction in Orthopedic Surgery: A Comparative Evaluation

View Full Paper
AWAshton WilliamsonNTNazgol TavabiNKNishita Kalepalli

Key Points

  • LLMs outperformed administrative coding, improving extraction accuracy of surgical procedures, which is essential for patient care.
  • Achieving macro-F1 scores above 0.6, the models demonstrated significant gains, with larger models further boosting performance.
  • Evaluation conducted on 800 clinical notes highlighted strengths and weaknesses in LLMs for less common procedures.
  • Future applications point toward faster and cheaper registry maintenance, though aligning fully with surgical experts remains a challenge.

Abstract

Background Operative notes in electronic health records contain critical information for understanding surgical care, yet manual coding is time-consuming, costly, and inconsistent. Large language models (LLMs) promise to transform this process by automatically extracting detailed procedure information—a capability with significant implications for scaling clinical registries and advancing surgical research. Methods We conducted a large-scale evaluation of state-of-the-art LLMs for zero-shot structured information extraction from orthopedic clinical notes. Fourteen open-source and proprietary models were tested on 800 real operative notes, annotated by both an orthopedic surgeon and an administrator using a curated list of 74 procedure classes. We compared model outputs to human annotations, assessing accuracy and exploring the effects of model scale, reasoning capabilities, and prompt design. Results Across models, LLMs consistently outperformed administrator-assigned labels, achieving macro-F1 scores above 0.6 and improving over administrative coding by up to 10 points. Larger models and reasoning capabilities further boosted performance, though gains plateaued beyond 30 billion parameters. Performance varied by procedure frequency, revealing clear strengths and persistent challenges for rare or complex cases. Conclusion Modern LLMs can already outperform routine administrative coding in extracting detailed surgical procedure data, pointing to a future where registry curation could be faster, cheaper, and more consistent. Yet, full alignment with surgical experts remains an open challenge—especially for rare procedures—emphasizing the need for domain adaptation and thoughtful deployment. Our findings illustrate how general-purpose LLMs can advance automated clinical data curation and inform the next generation of surgical informatics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Williamson et al. (2025) studied this question.

synapsesocial.com/papers/68af5bc7ad7bf08b1eae024fhttps://doi.org/10.1101/2025.08.19.25333995
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Comparison of Large Language Models with Rules-Based Natural Language Processing Algorithms for Extracting Data from Operative Notes2026
  2. 2Large Language Models in Surgery: Promise, Pitfalls, and Practical Use2026 · 1 citations
  3. 3Large Language Models in Spine Surgery : A Narrative Review of Performance Paradox and Clinical Integration Challenges2026
  4. 4Large Language Models Improve Operative Note Coding Accuracy and Financial Outcomes in Neurotology2026
  5. 5Large language models in laparoscopic surgery: A transformative opportunity2024 · 2 citations