PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 19, 2026Sensors0 citationsOpen Access

JMSC: Joint Spatial–Temporal Modeling with Semantic Completion for Audio–Visual Learning

View Full Paper
XXXinfu XuFYFan YangZYZhibin Yu

Key Points

  • This work aims to improve the integration of audio and visual information in understanding dynamic scenes.
  • Proposed JMSC framework for joint spatial-temporal modeling
  • Introduced cross-modal latent reconstruction for enhanced semantic completion
  • Leveraged audio guidance for unified representation of spatial and temporal attributes
  • Achieved state-of-the-art performance across various tasks
  • Demonstrated high computational efficiency compared to prior methods
  • Enhanced understanding of dynamic video context through better integration of modalities

Abstract

‌Audio–visual learning‌ seeks to achieve holistic scene understanding by integrating auditory and visual cues. Early research focused on fully fine-tuning pre-trained models, incurring high computational costs. Consequently, recent studies have adopted ‌parameter-efficient tuning‌ methods to adapt large-scale vision models to the audio–visual domain. Despite the competitive performance of existing methods, several challenges persist. Firstly, effectively leveraging the ‌complementary semantics‌ between the audio and visual modalities remains difficult, as these two modalities capture fundamentally different aspects of a video. Secondly, comprehending ‌dynamic video context is challenging because both spatial attributes (such as scale) and temporal characteristics (such as motion) of objects co-evolve over time, making semantic comprehension more complex. To address these challenges, we propose a novel framework, named Joint Spatial–Temporal Modeling with Semantic Completion (JMSC). JMSC introduces cross-modal latent reconstruction, which moves beyond shallow correlation by encouraging the model to reconstruct one modality’s complete semantic summary from a masked version of its counterpart. Furthermore, JMSC learns a unified representation of video spatial attributes and temporal changes by jointly modeling them under audio guidance, enabling accurate localization and consistent tracking in dynamic video scenes. Experimental results demonstrate that JMSC achieves state-of-the-art performance across multiple downstream tasks while maintaining high computational efficiency.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xu et al. (2026) studied this question.

synapsesocial.com/papers/6996a7e3ecb39a600b3edf5bhttps://doi.org/10.3390/s26041288
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1SHMamba: Structured Hyperbolic State Space Model for Audio-Visual Question Answering2024
  2. 2CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering2024
  3. 3Semantic-Assisted Object Clustering for Multi-Modal Referring Video Segmentation2025 · 1 citations
  4. 4Visual Jigsaw Post-Training Improves MLLMs2025
  5. 5Unsupervised Audio-Visual Segmentation with Modality Alignment2024 · 1 citations