ChatGPT achieved 70% complete agreement with an expert pulmonologist for cardiopulmonary exercise test interpretation, with a case-level κ = 0.40 and domain-level κ = 0.55.
Cross-Sectional (n=10)
Single-blind
Does ChatGPT interpretation of cardiopulmonary exercise tests agree with board-certified pulmonologists?
ChatGPT demonstrated moderate agreement (70% concordance) with expert pulmonologists for CPET interpretation, suggesting potential utility as an educational or assistive tool.
Effect estimate: κ = 0.40
Abstract Rationale Artificial intelligence (AI) has demonstrated potential in automating the interpretation of physiological data, including pulmonary function testing. In contrast, cardiopulmonary exercise testing (CPET), which assesses how the lungs, heart, and muscles respond during exercise, is more complex and requires integrating ventilatory (breathing), cardiac (heart), and metabolic (energy use) domains that are often subject to interpretation. Large language models (LLMs), such as ChatGPT, have recently been investigated for clinical reasoning; however, their effectiveness in structured CPET interpretation remains uncharacterized. This pilot study assessed the level of agreement between ChatGPT and board-certified pulmonologists in interpreting CPETs and evaluated reproducibility through blinded adjudication. Methods Ten de-identified CPETs were analyzed. ChatGPT was prompted using a structured schema that included the following components: presence and severity of airflow obstruction, presence and severity of exercise limitation, achievement of maximum heart rate, status of respiratory reserve (preserved or reduced), physiological dead space (VD) to tidal volume (VT) ratio (resting value and change with exercise), attainment of anaerobic threshold, presence of oxygen desaturation and need for oxygen support, and mechanism of exercise limitation (respiratory, cardiac, both, or neither). A board-certified pulmonologist (Pulmonologist 1) independently interpreted each study while blinded to the AI results. Discordant cases were subsequently reviewed by a second board-certified pulmonologist (Pulmonologist 2), who received de-identified paired outputs (AI versus Pulmonologist 1) and was tasked with adjudicating agreement. Results Among the 10 studies, 7 (70%) demonstrated complete agreement between ChatGPT and Pulmonologist 1, while 3 (30%) were discordant. In discordant cases, category-level differences were observed in airflow obstruction (2/10, 20%) and mechanism of exercise limitation (2/10, 20%), with no differences in other categories. Adjudication by Pulmonologist 2 favored ChatGPT in 1 case (10%) and Pulmonologist 1 in 2 cases (20%). At the case level, agreement between ChatGPT and Pulmonologist 1 was κ = 0.40. When assessed across all interpretive domains (total paired datapoints = 80), agreement improved to κ = 0.55. Overall, the agreement between pulmonologists 1 and 2 after adjudication remained high (κ = 0.80). Conclusions This pilot analysis found ChatGPT achieved 70% concordance with an expert pulmonologist for CPET interpretation. Adjudication revealed ChatGPT was correct in one-third of discordant cases. These findings highlight both the current limitations and the distinct educational potential of open-access LLMs for complex physiological interpretation. Larger, methodologically robust studies are needed to establish whether ChatGPT can consistently assist trainees in guideline-based, reproducible integration of ventilatory, cardiac, and metabolic data. This abstract is funded by: None
Abbas et al. (Fri,) conducted a cross-sectional in Cardiopulmonary exercise testing (CPET) interpretation (n=10). ChatGPT vs. Board-certified pulmonologists was evaluated on Complete agreement between ChatGPT and Pulmonologist 1 (κ = 0.40). ChatGPT achieved 70% complete agreement with an expert pulmonologist for cardiopulmonary exercise test interpretation, with a case-level κ = 0.40 and domain-level κ = 0.55.