PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 28, 2026Electronics0 citationsOpen Access

RAFFNet: Restricted Attention Feature Fusion Network for Self-Supervised Image Representation Learning

View Full Paper
JLJianeng LiFCFei ChenSZShufen Zhang

Key Points

  • To improve image representation learning using a novel model that integrates multi-level features and spatial information.
  • Developed RAFFNet for feature fusion using self-supervised learning.
  • Implemented dual attention mechanism focusing on channel and spatial dimensions.
  • Created attentional weighted masks to balance local and global feature extraction.
  • RAFFNet outperformed existing image representation models across several datasets.
  • Achieved enhanced generalization and representation performance metrics on CIFAR-10, CIFAR-100, Tiny ImageNet, and ImageNet-1%.
  • Demonstrated improved performance on object detection tasks using PASCAL VOC and COCO.

Abstract

Learning image representations with deep self-supervised models is an important task in computer vision, which aims to establish beneficial and general representations from unlabeled images. However, existing efforts train models mainly on high-level features, neglecting lower-level features and their global spatial information, thus limiting the discriminative power of the learned representations. In this work, we propose a representational learning model based on restricted attention feature fusion network (RAFFNet) to improve the quality and generalization of the learned image representations. Specifically, to fully exploit the features in the deep network, we use a self-supervised model on multi-level features to learn more general representations. Meanwhile, a new feature fusion strategy with a dual attention mechanism of channel and space is used for multi-level features, enabling the model training to obtain more important and comprehensive feature information. Furthermore, in order to better extract global spatial information, we devise a simple but effective attentional weighted mask, which restricts the weight of spatial attention and prevents the model from focusing only on local features with high attention weights. Experiments on four public classification datasets, CIFAR-10, CIFAR-100, Tiny ImageNet and ImageNet-1%, and two object detection datasets, PASCAL VOC and COCO, demonstrate that the proposed RAFFNet has better representation performance and generalization ability than most state-of-the-art image representation learning algorithms.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/69a287240a974eb0d3c02a51https://doi.org/10.3390/electronics15050964
Ask AI
Helpful
Bookmark
Share
View Full Paper