Abstract Background Surveillance of surgical site infections (SSI) typically relies on manual methods that are time-consuming and prone to subjectivity. Automation using artificial intelligence (AI), machine learning, and large language models, could enhance efficiency and broaden the scope of surveillance. This study evaluates the diagnostic accuracy of ChatGPT in detecting SSI through the analysis of electronic health records (EHRs) following colorectal surgery, comparing its performance to that of a manual SSI surveillance program. Methods Retrospective cohort study of patients who underwent elective colorectal surgery between 2010 and 2023. SSI was defined according to CDC-NHSN and ECDC criteria: superficial incisional (SSI-S), deep incisional (SSI-D), and organ/space (SSI-O/S). A standardized prompt was developed for ChatGPT. Clinical notes from EHRs were automatically extracted for analysis by ChatGPT 4o and compared to results from the manual surveillance program. Univariate and bivariate analyses were conducted. Sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), AUC, and ROC curves were calculated. Results A total of 1,213 patients were included. Manual surveillance identified an SSI rate of 11.2% (SSI-S 3.1%, SSI-D 1.2%, SSI-O/S 6.8%). For overall SSI, ChatGPT showed sensitivity and specificity of 86%, with PPV 44%, NPV 99%, and AUC 0.86 (80.6–89.9). No significant differences were observed between colon and rectal surgeries. Results for SSI-O/S were similar: sensitivity 86%, specificity 85%, PPV 99%, NPV 44%, and AUC 0.86 (80.6–89.9). However, diagnostic performance for SSI-S and SSI-D was substantially lower, with sensitivities of 21% and 30%, specificities of 98% and 99%, PPVs of 97% and 99%, NPVs of 44% and 44%, and AUCs of 0.6 (49.4–70.3) and 0.6 (43.3–76.4), respectively. Conclusions ChatGPT demonstrated acceptable sensitivity and specificity for detecting overall SSI and SSI-O/S. The high NPV allows for reliable exclusion of non-infected cases, while the low PPV indicates the presence of false positives. ChatGPT may serve as a useful tool for semi-automated surveillance and workload reduction in manual review processes. The low PPV highlights the need for further refinement. Integration of ChatGPT-based EHR analysis with additional data on antibiotic treatment, imaging, and microbiological test results could enhance SSI surveillance.
Badia et al. (2026) studied this question.