Single-cell gene expression profiling has emerged as a powerful technology for dissecting complex tissues at unprecedented resolution. Accurate cell clustering is a fundamental computational prerequisite for cell type identification. In recent years, numerous single-cell contrastive clustering algorithms have been developed to more effectively identify individual cells and characterize cellular heterogeneity. Despite these advances, the presence of noise and data sparsity in single-cell datasets continues to impede algorithmic performance. Mainstream clustering methods face three principal limitations: (i) most clustering methods process data in mini-batches, and batch-related value shifts can destabilize feature extraction; (ii) existing contrastive learning strategies fail to achieve multilevel, coordinated information fusion, making them prone to collapsed solutions; and (iii) noise and sparsity can cause centroid shifts, yet existing methods lack effective mechanisms for dynamic centroid updating. To address these challenges, we propose sample-wise debiased multilevel contrastive clustering with shrinkage risk regularization (SDMCC). Instead of directly learning from the original sparse data, SDMCC applies a sample-wise correction module in conjunction with a batch-size adaptation module to mitigate distortion. Furthermore, we introduce a multilevel contrastive learning strategy that integrates instance-level and cluster-level information, enabling the capture of fine-grained cellular variations while maintaining coarse-grained semantic consistency across cell subpopulations. In addition, a shrinkage risk regularization term prevents convergence to trivial solutions. Extensive experiments conducted on multiple single-cell datasets demonstrate that SDMCC not only achieves superior clustering accuracy but also effectively uncovers biologically meaningful patterns in gene expression, thereby highlighting its potential for broader biomedical applications.
Han et al. (Thu,) studied this question.