Abstract Rationale Artificial intelligence (AI) systems have demonstrated high diagnostic accuracy in interpreting pulmonary function tests (PFTs) when using proprietary algorithms. However, these closed-source tools are inaccessible for educational purposes or independent validation. Open-access large language models (LLMs), including ChatGPT, have recently been investigated for medical reasoning tasks but remain largely untested in physiological education. This pilot study evaluated the concordance between ChatGPT and board-certified pulmonologists in interpreting complete PFTs and assessed reproducibility through blinded adjudication. Methods Ten de-identified PFTs, including spirometry, lung volumes, and diffusing capacity, were analyzed. ChatGPT was prompted with standardized 2022 American Thoracic Society (ATS) and European Respiratory Society (ERS)3 interpretive rules to generate structured one-line summaries across five categories: airflow obstruction (presence/absence and severity), bronchodilator response (presence/absence), restriction (presence/absence and severity), hyperinflation or air trapping (presence/absence), and diffusing capacity of the lung for carbon monoxide (DLCO), including severity and correction for alveolar volume. A board-certified pulmonologist (Pulmonologist 1) independently interpreted each study while blinded to the AI results. For discordant cases, a second pulmonologist (Pulmonologist 2) reviewed de-identified paired outputs and selected the interpretation considered correct. Results Among the 10 studies, 4 (40%) demonstrated full agreement between ChatGPT and Pulmonologist 1, while 6 (60%) were discordant. Category-level discordances included airflow obstruction in 3 of 10 cases (30%), bronchodilator response in 1 of 9 cases (11%), restriction in 3 of 10 cases (30%), hyperinflation or air trapping in 1 of 10 cases (10%), and DLCO interpretation in 2 of 10 cases (20%). In all 10 category-level discordances, Pulmonologist 2 independently agreed with Pulmonologist 1 in every instance, confirming that all discrepancies were due to incorrect AI interpretations. Pulmonologist 1 demonstrated perfect agreement (κ = 1.0), whereas ChatGPT’s agreement with Pulmonologist 2 was κ = 0.25 at the case level and κ = 0.45 across all interpretive domains. Conclusions In this pilot analysis, ChatGPT produced concordant interpretations in 40% of PFTs. All discordant results were adjudicated in favor of the board-certified pulmonologist, highlighting the limitations of current open-access LLMs for clinical interpretation. However, structured prompting enabled reproducible, guideline-based outputs that may serve as educational scaffolds for trainees learning ATS/ERS-aligned PFT interpretation. Further research with larger datasets and refined prompts is warranted to better define ChatGPT’s role in physiological education. This abstract is funded by: None
Abbas et al. (Fri,) studied this question.