PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 1, 2026CAAI Transactions on Intelligence Technology0 citationsOpen Access

Video Super‐Resolution via Effective Spatio‐Temporal Alignment Network

View Full Paper
BGBin GuoXWXin WangHWHao Wen

Key Points

  • The aim is to develop a new network, ESTA-Net, to improve spatio-temporal alignment in video super-resolution.
  • Proposed an alignment module utilizing cascaded group convolutions to estimate motion offsets.
  • Implemented a bi-scale alignment strategy to manage complex motion effectively.
  • Introduced an attention-based feature enhancement module to refine aligned features.
  • ESTA-Net outperforms existing video super-resolution methods on standard benchmarks.
  • Reduced computational cost while achieving a wider receptive field for alignment accuracy.
  • Demonstrated a balance between model size and performance, indicating its practical applicability.

Abstract

ABSTRACT Extracting spatio‐temporal cues from neighbouring frames is challenging in video super‐resolution (VSR). Although deformable alignment‐based VSR methods have shown promise in aligning neighbouring frames with the reference frame, most existing methods rely on one or a few traditional convolutions to estimate motion offsets for spatio‐temporal alignment, restricting receptive field size and alignment accuracy. To address these limitations, we propose an effective spatio‐temporal alignment network (ESTA‐Net) for VSR. The core component of our method is the group convolution‐based alignment module (GCBAM), which utilises cascaded group convolutions to learn offsets across both the original and downsampled resolutions. By employing group convolutions rather than traditional convolutions, GCBAM enables the deformable alignment to achieve a wider receptive field with lower computational cost, thereby improving the accuracy of offset estimation. Additionally, the bi‐scale alignment strategy within GCBAM enhances robustness to complex and large‐scale motions. Furthermore, we introduce an attention‐based feature enhancement module (AFEM) to refine the aligned features, focusing on critical details to improve reconstruction quality. Extensive experiments on standard benchmarks show that our ESTA‐Net achieves superior VSR performance against other advanced methods, while maintaining a good equilibrium between model size and performance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Guo et al. (2026) studied this question.

synapsesocial.com/papers/6a1d234302fbce9130638e35https://doi.org/10.1049/cit2.70151
Ask AI
Helpful
Bookmark
Share
View Full Paper