A Vision Transformer-based framework (RegionViT) accurately performed automated COPD staging (balanced accuracy 93%; 95% CI 91-95%) and emphysema grading (88%; 95% CI 84-92%) using CT imaging.
Observational (n=3,840)
Does a Vision Transformer (ViT)-based framework accurately perform automated COPD staging and emphysema grading using CT imaging?
A Vision Transformer-based deep learning model demonstrated high accuracy and precision for automated COPD staging and emphysema grading from CT scans.
Effect estimate: Balanced accuracy 93% (COPD) and 88% (emphysema) (95% CI 91-95% (COPD), 84-92% (emphysema))
Abstract Rationale Chronic obstructive pulmonary disease (COPD) is a major global health burden, frequently coexisting with emphysema and contributing to excess mortality. Although emphysema can be identified in computed tomography (CT) scans, COPD is primarily diagnosed using spirometric measures, although spirometry is not widely adopted and opportunistic screening of COPD based on CT imaging can be advantageous. Recent studies have applied artificial intelligence (AI) models for COPD and emphysema assessment, yet most rely on convolutional architectures that may not fully capture global contextual features. This study evaluates the performance of a Vision Transformer (ViT)-based framework for automated COPD staging and emphysema grading using CT imaging. Methods We analyzed 3,840 CT scans from the COPDGene Phase 1 cohort. Emphysema grade was defined by percent low attenuation area (LAA-950) as: none (5%), mild (5-10%), moderate (10-20%), and severe (20%). COPD severity included PRISm (−1), no COPD (0), and Global Initiative for Chronic Obstructive Lung Disease (GOLD) 1-4 stages. Each CT was subsampled into 20 equally spaced axial 2D slices across the lung region. RegionViT was trained on a subset of lung-centered 9 slices in two separate tasks for emphysema and COPD while assuming slice independence with a balanced label 80:20 training:validation split. Classification for each patient was obtained using a 2.5D multi-slice Bayesian probability fusion of all 20 slices. The primary metrics of performance with 95% low-high confidence intervals included linear-weighted Cohen’s kappa, balanced accuracy, and macro AUC. One-off metrics were also reported. Results In the validation set, emphysema grading with all 20 slices achieved a balanced accuracy=88% 84-92%, macro AUC=0.976 0.968-0.982, κ=0.86 0.82-0.89, one-off accuracy=94% 93-96%, and one-off κ=0.99 0.97-1.00.COPD staging yielded a balanced accuracy=93% 91-95%, macro AUC=0.990 0.987-0.993, κ=0.88 0.86-0.91, one-off accuracy=93% 91-94%, and one-off κ=0.92 0.89-0.94. Conclusions We propose the use of RegionViT, specifically its adapted 2.5D multi-slice implementation with Bayesian probability fusion as opportunistic screening of COPD GOLD staging and emphysema grading from CT scans. To the best of our knowledge, this is the first deep learning-based approach capable of successfully distinguishing between PRISm, healthy controls, and COPD severity stages with both accuracy and precision. The algorithm presented in this study surpasses prior models for COPD and emphysema diagnosis from CT scans, including those based on Elastic Net, convolutional, or residual neural networks. Further validation and external replication are needed to confirm these results. This abstract is funded by: This work was supported by NHLBI grants 1R01HL149877, U01 HL089897, and U01 HL089856 and by NIH contract 75N92023D00011
Martín-Saladich et al. (Fri,) conducted a observational in Chronic obstructive pulmonary disease (COPD) and emphysema (n=3,840). RegionViT (Vision Transformer-based framework) was evaluated on Automated COPD staging and emphysema grading performance (balanced accuracy, macro AUC, linear-weighted Cohen's kappa) (Balanced accuracy 93% (COPD) and 88% (emphysema), 95% CI 91-95% (COPD), 84-92% (emphysema)). A Vision Transformer-based framework (RegionViT) accurately performed automated COPD staging (balanced accuracy 93%; 95% CI 91-95%) and emphysema grading (88%; 95% CI 84-92%) using CT imaging.