A Vision Transformer framework (augViT) predicted 5-year lung-function trajectories from baseline CT scans in COPD patients with a macro-AUC of 0.82 and one-off accuracy of 83%.
Observational (n=3,840)
Can a Vision Transformer (augViT) accurately predict 5-year lung-function trajectories from baseline chest CT scans in patients with COPD?
Vision Transformers can impute 5-year lung-function trajectories from baseline chest CT scans in COPD patients, potentially expanding access to trajectory-based risk stratification without requiring serial spirometry.
Effect estimate: macro-AUC 0.82
Abstract Rationale Lung function trajectories summarize long-term patterns of growth and decline that are associated with respiratory morbidity and premature mortality. Despite their value as epidemiologic tools, trajectory assignment still relies on longitudinal spirometric data rarely available in routine care. Chest CT scans, widely acquired for clinical and research purposes, capture structural correlates of airflow limitation that could enable trajectory imputation at the individual level. Vision Transformers (ViTs), which capture global contextual information beyond convolutional architectures, offer new opportunities to infer such latent functional trajectories from imaging. We evaluated a ViT-based approach for predicting 5-year lung-function trajectories from baseline CT scans in COPD. Methods A 2.5D ViT framework (augViT) was developed to classify spirometry-defined lung-function trajectories using baseline CT scans. Each scan was equispatially subsampled into 20 axial slices; 9 central slices were processed with a RegionViT model trained using an 80:20 split and a cross-entropy loss. RegionViT features were extracted for each slice, with data augmentation applied. For each patient, augViT selected a random feature vector representing a single slice and fed it into a transformer, which performed multi-slice fusion to produce subject-level predictions using 9 and 20 slices. Model performance was evaluated using linear-weighted Cohen’s κ, balanced accuracy, macro-AUC, and one-off metrics. Results The study included 3,840 participants from COPDGene Phase 1 and 2 with paired CT and spirometry data at baseline and 5-year follow-up. Using augViT, participants were assigned to one of six previously defined trajectories (1: supranormal, n = 431; 2: normal, n = 931, 3-6: progressively impaired, n = 2,477). In Phase 1 validation using 20 slices, augViT achieved a balanced accuracy of 50% (κ = 0.46), macro-AUC=0.82, one-off accuracy=83%, and one-off κ = 0.73. Replication in Phase 2 scans showed comparable one-off accuracy (78%) with macro-AUC=0.74 and κ = 0.37 (one-off κ = 0.64). When evaluating the model consistency over 5-years, the agreement between Phase 1 and Phase 2 predictions was 84% (one-off 96%). Conclusions ViTs - and specifically augViT - enable subject-level imputation of lung-function trajectories from standard chest CTs, bridging structural imaging with longitudinal physiology. This approach could expand access to trajectory-based risk stratification without requiring serial spirometry, supporting precision monitoring and early identification of high-risk COPD subtypes. Further validation and external replication are warranted to assess the clinical utility of this approach. This abstract is funded by: This work was supported by NHLBI grants 1R01HL149877, U01 HL089897, and U01 HL089856 and by NIH contract 75N92023D00011
Martín-Saladich et al. (2026) conducted an observational in COPD (n=3,840). 2.5D ViT framework (augViT) was evaluated on Prediction of 5-year lung-function trajectories (macro-AUC 0.82). A Vision Transformer framework (augViT) predicted 5-year lung-function trajectories from baseline CT scans in COPD patients with a macro-AUC of 0.82 and one-off accuracy of 83%.