ConTEXTual Net 3D: Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings | Synapse