Hyperspectral image (HSI) classification remains challenging due to the high dimensionality of spectral data, the strong coupling between spatial and spectral information, and the difficulty of learning robust representations under limited annotated samples. To address this, we propose 3D W-Net, a novel deep learning framework that integrates 3D convolution into the U-Net architecture and introduces a parallel–serial multi-scale feature learning strategy tailored for HSIs. The encoder incorporates a parallel spectraldimension multi-scale feature processing module, which employs three 3D convolution kernels with different spectral receptive fields to capture diverse spectral-scale features. An adjacent-scale selective fusion strategy is then applied to enrich spectral representations while reducing redundancy. The decoder features a serial spatial-dimension multiscale feature restoration module that progressively merges high-level semantic and lowlevel spatial details, preserving spatial–spectral correlations and mitigating information loss typically caused by flattening operations. This design forms a distinctive “W”-shaped feature propagation path that enhances multi-scale feature interaction. Extensive experiments on five public datasets(Botswana, Pavia University, Chikusei, Houston 2013, and WHU-Hi-HongHu) demonstrate that 3D W-Net achieves superior performance compared to six state-of-the-art methods, attaining overall accuracies of 99.6%, 99.1%, 100%, 99.3%, and 99.5%, respectively. The results highlight the effectiveness of the proposed parallel–serial multi-scale strategy in improving classification accuracy and generalization capability for complex HSI scenarios.
Wang et al. (Thu,) studied this question.