PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20250 citationsOpen Access

MM-FusionNet: Context-Aware Dynamic Fusion for Multi-modal Fake News Detection with Large Vision-Language Models

View Full Paper
JHJunhao HeTLTianyu LiuJZJingyuan Zhao

Key Points

  • MM-FusionNet achieved a state-of-the-art F1-score of 0.938, demonstrating superior accuracy in multi-modal fake news detection.
  • The context-aware dynamic fusion module effectively adapts the importance of textual and visual features, improving detection performance.
  • Evaluation on the Multi-modal Fake News Dataset, comprising 80,000 samples, underscores the robustness of the model against modality perturbations.
  • Results highlight that MM-FusionNet's performance approaches human-level accuracy, suggesting strong practical applicability.

Abstract

The proliferation of multi-modal fake news on social media poses a significant threat to public trust and social stability. Traditional detection methods, primarily text-based, often fall short due to the deceptive interplay between misleading text and images. While Large Vision-Language Models (LVLMs) offer promising avenues for multi-modal understanding, effectively fusing diverse modal information, especially when their importance is imbalanced or contradictory, remains a critical challenge. This paper introduces MM-FusionNet, an innovative framework leveraging LVLMs for robust multi-modal fake news detection. Our core contribution is the Context-Aware Dynamic Fusion Module (CADFM), which employs bi-directional cross-modal attention and a novel dynamic modal gating network. This mechanism adaptively learns and assigns importance weights to textual and visual features based on their contextual relevance, enabling intelligent prioritization of information. Evaluated on the large-scale Multi-modal Fake News Dataset (LMFND) comprising 80,000 samples, MM-FusionNet achieves a state-of-the-art F1-score of 0.938, surpassing existing multi-modal baselines by approximately 0.5% and significantly outperforming single-modal approaches. Further analysis demonstrates the model's dynamic weighting capabilities, its robustness to modality perturbations, and performance remarkably close to human-level, underscoring its practical efficacy and interpretability for real-world fake news detection.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

He et al. (2025) studied this question.

synapsesocial.com/papers/68f0f51d8dd8ea469b1d705ehttps://doi.org/10.48550/arxiv.2508.05684
Ask AI
Helpful
Bookmark
Share
View Full Paper