PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 6, 2026International Journal of Advanced Computer Science and Applications0 citationsOpen Access

Enhancing Misinformation Detection on Twitter with a Content-Based Multi-Lingual Bert Model

View Full Paper
KKKrishna KumarAVAkila Venkatesan

Key Points

  • The goal is to improve real-time detection of misinformation on social media platforms using a multilingual model.
  • Developed a Content-based Attention Multi-lingual BERT (CA-BERT) model.
  • Utilized a balanced multilingual tweet dataset focused on COVID-19 topics.
  • Employed the LIME interpretability method for transparent prediction explanations.
  • Evaluated performance against baseline models like RoBERTa and DANN.
  • Achieved 96% recall for true information and 95% for misinformation in English.
  • Obtained F1 Scores of 93% for true information and 92% for misinformation.
  • Showed significant cross-lingual generalization, particularly for Dutch (75% F1) and Spanish (72% F1).
  • Identified performance challenges with Arabic due to tokenization issues.

Abstract

The rapid spread of misinformation during global crises like COVID-19 has severely impacted public health, governance, and social trust. Social media platforms such as Twitter have amplified this issue, underscoring the urgent need for multilingual, real-time misinformation detection. The proposed Content-based Attention Multi-lingual BERT (CA-BERT) model addresses this challenge by enhancing the standard BERT framework with a content-based attention mechanism that assigns adaptive weights to semantically important tokens often linked to false or misleading content. This attention enables deeper contextual understanding of misinformation cues across diverse linguistic contexts. Using the LIME interpretability method, CA-BERT provides transparent explanations of its predictions, supporting accountable decision-making for policymakers and content moderators. Leveraging multilingual BERT (mBERT) allows the model to handle multiple languages simultaneously, ensuring robust cross-lingual applicability. Evaluations using a balanced multilingual tweet dataset on COVID-19 topics demonstrate that CA-BERT outperforms baseline models such as RoBERTa, DANN, and HANN, achieving 96% recall for true information and 95% for misinformation in English, with F1 Scores of 93% and 92%, respectively. The model maintains strong cross-lingual generalization, especially for Dutch (75% F1) and Spanish (72% F1), with slightly lower performance for Arabic due to tokenization and dialectal complexity. These results highlight CA-BERT’s adaptability while underscoring the need for improved handling of low-resource, morphologically rich languages. Future work involves region-specific preprocessing, cross-lingual transfer learning, and multimodal misinformation detection, aiming to transform CA-BERT into a core component of multilingual real-time disinformation monitoring systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kumar et al. (2026) studied this question.

synapsesocial.com/papers/698586118f7c464f23009e27https://doi.org/10.14569/ijacsa.2026.0170169
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Bilingual COVID-19 Fake News Detection Based on LDA Topic Modeling and BERT Transformer2023 · 13 citations
  2. 2Detecting the Impact of COVID-19 on Social Media using BERT-Based Model2024
  3. 3PolyTruth: Multilingual Disinformation Detection using Transformer-Based Language Models2025
  4. 4BERT for Twitter Sentiment Analysis: Achieving High Accuracy and Balanced Performance2024 · 4 citations
  5. 5Leveraging the BERT Model for Enhanced Sentiment Analysis in Multicontextual Social Media Content2024 · 13 citations