PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 31, 2026Sensors0 citationsOpen Access

QA2FDet: Quality-Aware Adaptive Alignment Fusion Network for UAV RGBT Tiny Pedestrian Detection

View Full Paper
YTYifang TanLYLijun YuanCXChuanjiang Xie

Key Points

  • This research aims to enhance tiny pedestrian detection by addressing issues related to cross-modal misalignment and background interference in UAV imagery.
  • Developed QA2FDet, featuring spectrum-spatial decoupled enhancement, cross-modal correspondence mining, and prior-informed gated fusion.
  • Implemented discrete cosine transform for background disentanglement and deep semantic gating for noise suppression.
  • Utilized thermal-guided local asymmetric cross-attention for refining correspondences under slight spatial offsets.
  • QA2FDet achieved state-of-the-art performance on UAV RGBT detection benchmarks.
  • Demonstrated strong robustness in detecting tiny pedestrians even in challenging aerial scenes.

Abstract

Visible–thermal tiny pedestrian detection in UAV aerial images is crucial for online decision-making in urban security and disaster response. However, the extremely small scale and sparse distribution of pedestrians cause discriminative cues to be submerged by dominant low-frequency background and contextual redundancy during feature learning. Meanwhile, cross-modal spatial misalignment and spatially varying modality reliability hinder stable fine-grained correspondence, thereby degrading fusion quality. To address these issues, QA2FDet is proposed as a quality-aware adaptive alignment fusion network comprising three modules: spectrum-spatial decoupled enhancement module (SDE), cross-modal correspondence mining module (CCM), and prior-informed gated fusion (PGF). SDE leverages the discrete cosine transform to disentangle redundant low-frequency background information, while deep semantic gating propagates high signal-to-noise ratio details into shallow representations to enhance subtle cues of tiny pedestrians and suppress high-frequency noise. To establish fine-grained neighborhood correspondences under slight spatial offsets, thermal-guided local asymmetric cross-attention is designed in CCM. Finally, region-level quality and modality discrepancy are jointly modeled for adaptive cross-modal fusion in PGF. Extensive experiments on multiple UAV-based RGBT detection benchmarks demonstrate that QA2FDet achieves state-of-the-art performance and exhibits strong robustness in challenging aerial scenes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tan et al. (2026) studied this question.

synapsesocial.com/papers/6a1bd1745783ba022b6fcf91https://doi.org/10.3390/s26113443
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1RGB-T tracking by modality difference reduction and feature re-selection2022 · 24 citations
  2. 2Embedded Real-Time Object Detection for a UAV Warning System2017 · 117 citations
  3. 3Survey on computation offloading in UAV-Enabled mobile edge computing2022 · 225 citations
  4. 4TFDet: Target-Aware Fusion for RGB-T Pedestrian Detection2024 · 83 citations
  5. 5MultiSpectral Transformer Fusion via exploiting similarity and complementarity for robust pedestrian detection2025 · 23 citations