PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 5, 2026Inventions0 citationsOpen Access

Image Captioning Using Enhanced Cross-Modal Attention with Multi-Scale Aggregation for Social Hotspot and Public Opinion Monitoring

View Full Paper
SJShan JiangYCY. ChenRCRilige CHAOMU

Key Points

  • The central aim is to improve image captioning models for monitoring social hotspots and public opinion on social media.
  • Developed ECMA, a module for enhanced cross-modal interaction in image captioning.
  • Integrated ECMA into BLIP-2's Querying Transformer without altering the original architecture.
  • Implemented a multi-scale visual aggregation strategy and a semantic residual gating mechanism.
  • Increased CIDEr score from 144.6 to 146.8, a 1.52% improvement.
  • Improved BLEU-4 score from 42.5 to 43.9, a 3.29% increase.
  • Demonstrated superior coherence and informativeness in captions for complex social media images.

Abstract

Large volumes of images shared on social media have made image captioning an important tool for social hotspot identification and public opinion monitoring, where accurate visual–language alignment is essential for reliable analysis. However, existing image captioning models based on BLIP-2 (Bootstrapped Language–Image Pre-training) often struggle with complex, context-rich, and socially meaningful images in real-world social media scenarios, mainly due to insufficient cross-modal interaction, redundant visual token representations, and an inadequate ability to capture multi-scale semantic cues. As a result, the generated captions tend to be incomplete or less informative. To address these limitations, this paper proposes ECMA (Enhanced Cross-Modal Attention), a lightweight module integrated into the Querying Transformer (Q-Former) of BLIP-2. ECMA enhances cross-modal interaction through bidirectional attention between visual features and query tokens, enabling more effective information exchange, while a multi-scale visual aggregation strategy is introduced to model semantic representations at different levels of abstraction. In addition, a semantic residual gating mechanism is designed to suppress redundant information while preserving task-relevant features. ECMA can be seamlessly incorporated into BLIP-2 without modifying the original architecture or fine-tuning the vision encoder or the large language model, and is fully compatible with OPT (Open Pre-trained Transformer)-based variants. Experimental results on the COCO (Common Objects in Context) benchmark demonstrate consistent performance improvements, where ECMA improves the CIDEr (Consensus-based Image Description Evaluation) score from 144.6 to 146.8 and the BLEU-4 score from 42.5 to 43.9 on the OPT-6.7B model, corresponding to relative gains of 1.52% and 3.29%, respectively, while also achieving competitive METEOR (Metric for Evaluation of Translation with Explicit Ordering) scores. Further evaluations on social media datasets show that ECMA generates more coherent, context-aware, and socially informative captions, particularly for images involving complex interactions and socially meaningful scenes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jiang et al. (2026) studied this question.

synapsesocial.com/papers/69843564f1d9ada3c1fb41bdhttps://doi.org/10.3390/inventions11010013
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Image Captioning Model Based on Multi-Step Cross-Attention Cross-Modal Alignment and External Commonsense Knowledge Augmentation2025 · 5 citations
  2. 2Double‐Attention Transformer for Cross‐Modal Image Captioning: Enhancing Visual–Linguistic Alignment on Low‐Resource Datasets2026 · 1 citations
  3. 3Unpaired Image Captioning via Cross-Modal Semantic Alignment2025
  4. 4End‐to‐End Attention‐Enhanced Transformer for Image Captioning in Biomimetic Wearable Devices2025
  5. 5Context-aware image captioning using cosine transformer attention and optimal vision transformer with improved walrus optimization algorithm2026