PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 4, 2026Scientific Reports0 citationsOpen Access

Hand gesture 3D pose estimation method based on swin transformer and CNN

View Full Paper
RDRong DangGFGang Feng

Key Points

  • The aim is to improve gesture pose estimation accuracy by addressing feature extraction and joint relationships.
  • Utilized depth images as input for gesture feature extraction.
  • Implemented a convolutional network for initial feature extraction.
  • Adopted a Swin Transformer to capture long-range joint relationships and global features.
  • Employed a U-shaped network for hierarchical feature processing and preservation of joint information.
  • Introduced 2D Gaussian heatmaps for representing keypoint distributions during feature regression.
  • Achieved an average squared error reduction from 7.012 mm to 4.776 mm compared to the baseline model.
  • Demonstrated improved performance over state-of-the-art pose estimation networks.

Abstract

Existing gesture pose estimation methods commonly exhibit limitations such as singular feature extraction for hand characteristics and the neglect of long-range topological relationships between joints, thus restricting their prediction accuracy. To address these issues, this study proposes a gesture pose estimation method that takes depth images as input. First, rough gesture features are extracted using a convolutional network, while a Swin Transformer module captures the topological relationships between joints and global features through spatial information. A U-shaped network is used to process the features hierarchically, preserving local joint information at various resolutions, which is fused with global features. Two-dimensional Gaussian heatmaps are introduced to represent the distribution of keypoints, improving network supervision for target feature regression. The backend network outputs the final keypoint coordinates. We evaluated the method on a newly constructed dataset, achieving a reduced average squared error of 7.012–4.776 mm lower than that of the baseline model. The experimental results demonstrate the superior performance of the proposed method, when comparing it to state-of-the-art mainstream pose estimation networks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Dang et al. (2026) studied this question.

synapsesocial.com/papers/69a7ccf7d48f933b5eed8f5ahttps://doi.org/10.1038/s41598-026-41974-6
Ask AI
Helpful
Bookmark
Share
View Full Paper