ABSTRACT In 3D medical imaging, achieving accurate segmentation using a deep learning model is a vital task, but it is also important to understand how the models produce these results. In the deep learning model, they mostly get high performance, but their inner workings are difficult to understand. The healthcare sector is basically focused on accuracy, and something is missed, such as interpretability and model bias. Mostly, explanation models are designed for 2D data; when they are used in 3D data, they face hurdles in handling its complexity. This paper uses the voxel‐level attribution frameworks to focus on which parts of a 3D image are most influential in the model's prediction, and this can be done by using a global binary mask to highlight the most relevant regions and filter out the less important ones. The proposed frameworks use the model‐agnostic tool KernelSHAP, which makes the grouping in super‐voxel, which reduces the computational load without compromising explanation quality. This combined approach makes it easier to understand how the model works in complex medical scenarios. It provides clear and localized insight into the model decision. This framework supports more transparent and clinically trustworthy applications of deep learning in 3D medical image analysis.
Srivastava et al. (2026) studied this question.