PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 3, 2026Scientific Reports0 citationsOpen Access

A multi-scale feature fusion gaze estimation model based on convolutional neural network and vision transformer

View Full Paper
PWPeng WangXWX F WangSYShuo Yuan

Key Points

  • This study aims to improve gaze estimation accuracy through an innovative multi-scale feature fusion model called CAF-ViT.
  • Developed the CAF-ViT model to utilize multi-scale face images for gaze estimation.
  • Implemented a feature extraction process using ResNet-18 and introduced learnable Class Tokens.
  • Employed self-attention and cross-attention mechanisms to enhance feature representation.
  • Achieved gaze estimation errors of 3.73° on MPIIFaceGaze, an 8.6% improvement over the baseline.
  • Reported gaze estimation errors of 5.27° on EyeDiap and 10.43° on Gaze360.
  • Validation of the multi-scale fusion strategy through ablation studies.

Abstract

To address ineffective feature fusion and feature loss in gaze estimation under unconstrained environments, this study proposes a multi-scale feature fusion model, CAF-ViT (Cross-Attention Fusion Vision Transformer). The model takes multi-scale face images as input, uses ResNet-18 to extract feature maps at different granularities, and introduces learnable Class Tokens per scale. In the fusion stage, Class Tokens of different scales first perform self-attention in their respective Transformer Encoders to aggregate local details and global semantics. Then, by swapping the token sequences and computing cross-attention, the model achieves bidirectional interaction and deep fusion of coarse- and fine-grained features. To further refine feature representation, an additional attention layer is added after cross-attention. It linearly transforms the original query vector with Sigmoid activation to generate new query weights, and linearly maps the attention output to new value vectors, improving the representation of task-relevant features. The fused Class Token is finally regressed to gaze direction via a multilayer perceptron. Experiments show estimation errors of 3. 73^ on MPIIFaceGaze, representing an 8. 6% improvement over the baseline hybrid CNN-Transformer, and errors of 5. 27^ on EyeDiap and 10. 43^ on Gaze360. Ablation studies validate the multi-scale fusion strategy and improved attention mechanism.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2026) studied this question.

synapsesocial.com/papers/6a1fc530dee9eb8c0dce6a0ahttps://doi.org/10.1038/s41598-026-55923-w
Ask AI
Helpful
Bookmark
Share
View Full Paper