ABSTRACT This work aims to propose an algorithm to select interpretable and diverse patches from histopathology images, discarding irrelevant areas, for the computer‐aided diagnosis of breast and oral cancer. The proposed patch selection algorithm uses ResNet 50 features combined with Rotary Position Embedding for Vision Transformers with learned positional encoding. To evaluate patch interpretability and diversity, attention scores and cosine similarity are applied. This patch value metric identifies four key patches that contribute most to diagnosis while preserving diversity. Features from these selected patches are extracted using three pre‐trained networks, normalized via p ‐norm pooling, and classified through belief‐based fusion to ensure interpretability and diagnostic relevance in histopathological analysis. The proposed patch selection framework was evaluated on a publicly available breast and oral cancer histopathology datasets, achieving an accuracy of 98.46% and 97.84%, respectively, outperforming existing breast and oral cancer detection techniques. In addition, when evaluated on an independent oral cancer histopathology dataset using the model trained on the initial oral dataset, the framework maintained strong generalization capability. The proposed patch selection algorithm delivers improved diagnostic accuracy with optimized training and demonstrates strong generalization across both breast and oral cancer histopathology datasets. It provides an interpretable and effective framework for patch‐based classification, enhancing diagnostic reliability for pathologists and offering a robust solution for multi‐tissue cancer analysis.
Subhija et al. (Sun,) studied this question.