Background: Positron emission tomography (PET) provides physiologic information central to oncologic staging and treatment assessment, but its availability is limited by cost, radiation exposure, and scanner access. Synthesizing PET from computed tomography (CT) is attractive but challenging, as tracer uptake is only partially constrained by anatomy, making the mapping inherently one-to-many. Methods: We propose a conditional 3D latent diffusion framework (3D-LDM) for CT-to-PET synthesis in the head–neck and thoracic region. The pipeline localizes anatomy by segmenting lungs in CT and restricting the volume to reduce irrelevant variability. PET volumes are encoded into a compact latent space using a KL-regularized 3D autoencoder, and a conditional 3D diffusion U-Net learns to generate PET latents conditioned on CT via a denoising diffusion process. The model was trained and evaluated on 900 paired PET/CT studies. Performance was assessed in SUV space using MAE, PSNR, and SSIM, and compared against transformer-, CNN-, and GAN-based baselines. Results: On the held-out test cohort, 3D-LDM achieved the best overall quantitative fidelity (MAE = 303.05 ± 22.16 SUV units, PSNR = 32.64 ± 1.79, SSIM = 0.86 ± 0.03), outperforming all baselines with statistically significant differences (p < 0.001). At the lesion level, the model achieved a precision of 0.76 (95% CI: 0.71, 0.81) and recall of 0.76 (95% CI: 0.72, 0.80), detecting an average of 3.19 lesions per scan with a false-positive rate of 0.72/scan. Lesion-wise NMSE was 11.37%, significantly outperforming GAN and transformer baselines. Conclusions: 3D-LDM enables efficient, high-fidelity PET synthesis in the head–neck and thoracic regions, substantially improving lesion-level accuracy over state-of-the-art baselines. While it is not a replacement for diagnostic PET, these results support the model’s potential as a clinical decision support tool.
Mahdi et al. (2026) studied this question.