Abstract Background Pulmonary nodules are commonly encountered in lung cancer screening. The risk of malignancy varies widely and is generally estimated using expert consensus guidelines (Lung-RADS). Purpose To assess the performance of a deep learning algorithm (DeepPNP) for pulmonary nodule malignancy risk estimation in a lung cancer screening dataset and the effect of data enrichment in model training. Materials and Methods A retrospective analysis was conducted using three datasets. DeepPNP is a 3D convolutional network (EfficientNet-B0–based) operating on nodule-centered 3D patches. For the DeepPNP model training and validation, the National Lung Screening Trial (NLST) dataset was combined with two independent malignant nodule-only datasets, resulting in a merged dataset of 28,057 nodules, including 2,362 malignant nodules. An ablation model (DeepPNP-NLST) was trained on NLST only. The testing was conducted on a held-out data set from the NLST dataset. Performance metrics, including sensitivity, specificity, precision, F1 score, and accuracy, were analyzed across three operating thresholds selected based on specificities of 0.80, 0.85 and 0.90 (selected on the validation set). Benchmarks included Lung-RADS v2022 and the PanCan model. Results On the NLST test set (including 2,597 nodules from 1,243 CT scans, DeepPNP achieved an area under the receiver operating characteristic curve (ROC AUC) of 0.96 (95% CI: 0.95-0.97), outperforming Lung-RADS AUC = 0.91 (95% CI: 0.89-0.94; P.001) and PanCan AUC = 0.93 (95% CI: 0.91-0.95; P.001). DeepPNP-NLST had an AUC of 0.95 (95% CI: 0.93-0.97; P=.045 vs. DeepPNP), indicating a modest gain from positive-only supplementation. Subgroup analyses showed consistent outperformance across nodule sizes and types. Operating-point metrics at 0.80/0.85/0.90 specificity are reported; at 0.80 specificity, DeepPNP achieved sensitivity of 0.94 (100/107; 95% CI: 0.88–0.98) and specificity of 0.88 (2,196/2,490; 95% CI: 0.87–0.90). Conclusion DeepPNP outperformed established malignancy risk models in lung cancer screening. The inclusion of biopsy-confirmed malignant nodules from two external datasets provided a measurable performance gain, underscoring the importance of data enrichment during model training.
Barbosa et al. (Fri,) studied this question.