PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 14, 2026IEEE Transactions on Image Processing0 citations

Towards Universal Semantic Communication via Matchable Semantic Subspace Transmission

View Full Paper
BLBohan LiXYXi YangSDSongsong Duan

Key Points

  • The aim is to develop a Universal Semantic Communication framework that can communicate across unknown semantic categories under bandwidth constraints.
  • Proposed a Matchable Semantic Subspace Transmission framework for communication in arbitrary semantic categories.
  • Developed components including Visual Semantic Engine, Semantic Squeeze Network, Noise-Adaptive Re-expansion, and VLM-based Decoder.
  • Utilized a two-stage training strategy to optimize cross-modal alignment and transmission robustness.
  • Achieved state-of-the-art performance in semantic segmentation under low-SNR and extreme-compression conditions.
  • Demonstrated significant robustness and generalization compared to existing communication methods.

Abstract

Semantic communication targets reliable task execution at the receiver under stringent bandwidth and channel constraints. However, existing communication paradigms either focus on bit-level signal reconstruction, impeding the balance between task efficacy and bandwidth efficiency, or are limited by fixed vocabularies and lack generalization when facing unknown categories and open scenarios. To this end, we propose Universal Semantic Communication (UniSC), an open-vocabulary semantic communication framework that formulates transmission as a Matchable Semantic Subspace Transmission (MSST) problem. In this work, "universal" refers to the ability to handle arbitrary text-defined semantic categories beyond fixed vocabularies, rather than universality across all vision tasks. The transmitted representation is explicitly constrained to preserve cross-modal matchability after noisy transmission, rather than merely supporting latent recovery or closed-set inference. Concretely, UniSC comprises a Visual Semantic Engine (VSE), a Semantic Squeeze Network (SSN), a Noise-Adaptive Semantic Re-expansion (NASR) module, and a VLM-based Decoder. VSE and SSN project images into a compact semantic subspace for transmission. This subspace is optimized to preserve both robustness and cross-modal matchability under channel corruption. NASR denoises and lifts the received features back into a semantically complete visual space, from which the VLM-based Decoder performs open-category inference by matching arbitrary text queries rather than relying on a fixed classifier head. The VLM-based Decoder employs a Text Semantic Engine (TSE) to map natural language to text embeddings and, via a learnable Text-Visual Bridge (TVB), aligns them with the reconstructed visual structure for cross-modal matching. To improve cross-modal alignment and transmission robustness, a two-stage training strategy first establishes cross-modal anchors and then optimizes end-to-end robustness and compactness. Extensive experiments on semantic segmentation benchmarks demonstrate that UniSC achieves strong generalization and state-of-the-art performance under harsh channel conditions, outperforming existing methods in both low-SNR and extreme-compression regimes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/6a05659da550a87e60a1dfb3https://doi.org/10.1109/tip.2026.3690331
Ask AI
Helpful
Bookmark
Share
View Full Paper