Integrating spatial histology with multi-slide, multi-omics data is essential for deciphering tissue architecture and cellular dynamics at high resolution. However, incomplete modality overlap across sections hinders coherent integration and cross-condition analysis. Here, we present stMixer, an unsupervised framework that (i) employs self-looped cross-attention to jointly encode histological, molecular, and spatial features; (ii) implements a multi-modal metric learning module to achieve biologically coherent integration across sections; and (iii) uses a graph-guided, cluster-level voting algorithm to enable anatomically faithful label propagation. Benchmarking across six spatial modalities demonstrates that stMixer achieves superior scalability and accuracy in dimensionality reduction, batch correction, and label transfer. The framework accommodates large, heterogeneous datasets across tissues, species, and technologies. We further showcase its versatility in mosaic integration, pseudo-time inference, and cross-tissue knowledge transfer. Notably, stMixer uncovers transient thymic states overlooked by competing methods, resolves fine-grained cortical microstructures, and corrects anatomical mis-annotations through integration with single-cell reference. stMixer is available at https://github.com/YQX-code/stMixer/.
Yang et al. (Tue,) studied this question.