Abstract Introduction ARDS is under-recognized in clinical practice and patients do not receive evidence-based treatments. In research settings, identifying patients meeting ARDS criteria requires labor-intensive clinical adjudication. It is unknown whether large language models (LLMs) are capable of accurately reviewing clinical records to identify ARDS patients. Methods We studied consecutive patients with acute hypoxemic respiratory (PaO2/FiO2 300) receiving invasive ventilation, non-invasive ventilation, or heated high flow nasal cannula at a single center. All patients had previously been adjudicated for ARDS by multiple critical care physicians. We developed a structured LLM prompt using a pilot cohort of 96 patients, then evaluated performance on a temporally distinct validation cohort (Figure). The prompt instructed the LLM to determine whether a patient met ARDS criteria after reviewing relevant documentation (history and physical, progress notes, radiology reports, discharge summary, and calculated PaO2/FiO2 ratios). The prompt directed the LLM to provide a 9-point Likert scale confidence rating and item-level assessment for each ARDS criterion. Prompts were sent to a secure, private GPT-5.0 connection authorized for use with protected health information. LLM performance was compared to a majority vote physician reference standard. Results The LLM was tested on 416 patients, including 74 who met ARDS criteria, each reviewed by 2 to 8 critical care physicians. The LLM achieved an AUROC of 0.91 (95% CI 0.88-0.95) for ARDS identification. When the LLM rated ARDS at least “likely,” sensitivity was 80% (95% CI 69-88%), specificity 83% (95% CI 78-86%), positive predictive value 50% (95% CI 40-59%), and negative predictive value 95% (92-97%). Inter-rater reliability between physicians for identifying ARDS was kappa = 0.44. Inter-rater reliability between separate reviews by the LLM was kappa = 0.86. Inter-rater reliability between the LLM and physicians was a kappa = 0.50. LLM performance for each ARDS criterion was 1) timing within 1 week: 94% accuracy; 2) oxygenation: 88% accuracy; 3) bilateral airspace disease: 80% accuracy; 4) exclusion of cardiac edema: 74% accuracy. When asked to identify the time a patient first met ARDS criteria, the LLM estimated the onset a median of 2.4 hours (IQR 0.2-8.2) after the physician determined onset. Conclusions Large language models can review relevant clinical notes and identify patients meeting ARDS criteria with high accuracy and greater reliability than physicians. These findings support LLMs as a scalable tool to improve ARDS clinical recognition and enable high-fidelity cohort assembly for research and quality initiatives. This abstract is funded by: NIH
Sjoding et al. (Fri,) studied this question.