PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 16, 2026Alexandria Engineering Journal0 citationsOpen Access

Multi-scale adaptive small-object detection with edge fidelity and global semantic modeling

View Full Paper
QCQian CuiENEnhao NingMDMinghua Du

Key Points

  • The aim is to enhance small-object detection by solving detail retention, context modeling, and scale adaptivity challenges.
  • Developed a unified detection system combining adaptive cross-layer cooperation with multi-scale feature fusion.
  • Utilized a Hybrid Dual-Backbone Encoder integrating a Vision Transformer and deformable convolutions.
  • Implemented the Multi-Scale Attention Fusion Neck for similarity-guided feature aggregation.
  • The proposed framework significantly improved feature representation for small objects across various scenarios.
  • Extensive experiments indicated enhanced detection performance in small object recognition.
  • Ablation studies showcased the effectiveness of each component in improving accuracy.

Abstract

The combined challenges of detail retention, context modeling, and scale adaptivity continue to make small-object detection tough. Convolutional operators have trouble capturing long-range dependencies, attention-based representations may decrease already sparse object cues, and early downsampling frequently eliminates fine-grained structures. Furthermore, localization is more susceptible to representation bias and supervision is more brittle due to the lack of object pixels. We suggest a unified detection system that combines adaptive cross-layer cooperation with multi-scale feature fusion to solve these problems. The Hybrid Dual-Backbone Encoder (HDBE) combines a Vision Transformer for global contextual reasoning with anti-aliased deformable convolutions for local detail extraction. While Dilated-Window Self-Attention (DWSA) improves context interaction by expanding the effective receptive field, Token-Selective Downsampling (TSD) maintains informative representations during feature reduction. The Multi-Scale Attention Fusion Neck (MSAFN) performs similarity-guided feature aggregation across scales, and the Small-Object-Aware Decoupled Head (SADH) improves size-aware prediction and supervision. Extensive experiments and ablation studies demonstrate that the proposed framework effectively improves feature representation and detection performance for small objects across diverse scenarios.The Small-Object-Aware Decoupled Head (SADH) enhances size-aware prediction and supervision, while the Multi-Scale Attention Fusion Neck (MSAFN) carries out similarity-guided feature aggregation across scales. The suggested architecture successfully enhances feature representation and detection performance for small objects in a variety of settings, as shown by extensive experiments and ablation investigations.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cui et al. (2026) studied this question.

synapsesocial.com/papers/6a080b17a487c87a6a40d25chttps://doi.org/10.1016/j.aej.2026.04.056
Ask AI
Helpful
Bookmark
Share
View Full Paper