Vision-language foundation models have emerged as powerful general-purpose representation learners, but their deterministic embeddings often fail to provide the reliability required for high-stakes biomedical tasks. We introduce MedProbCLIP, a probabilistic vision-language learning framework for chest X-ray and radiology report representation learning and retrieval. Evaluated on the MIMIC-CXR dataset, MedProbCLIP outperforms deterministic and probabilistic baselines in both retrieval and zero-shot classification, while improving the trustworthiness significantly.
Elallaf et al. (Thu,) studied this question.