PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 2, 2026SHILAP Revista de lepidopterología0 citationsOpen Access

An Ensemble Method for Variable-Grain Political Ideology Classification on Social Media Texts

View Full Paper
EKErik-Robert KovacsFSFlorinel-Marian SgardeaAŞAurelia Ştefănescu

Key Points

  • This research aims to classify political ideology in social media texts using a novel ensemble deep learning approach.
  • Developed an ensemble deep learning classification pipeline for ideology categorization.
  • Trained models on a dataset from political news websites and tested on a sample of 2.6M tweets.
  • Used BERT classifiers, 5-fold cross-validation, and n-gram analysis for results validation.
  • Achieved an F1 score of 96.39% for political classification, 92.84% for coarse-grained ideology, and 90.33% for fine-grained ideology.
  • Utilized s-BERT embeddings and DBSCAN clustering to reveal relevant discussion topics.
  • SHAP explainers verified accuracy and identified stylistic features characteristic of each political orientation.

Abstract

During the early 21st century, the rise of social media has significantly affected all areas of social existence, including politics. Among other phenomena, this resulted in the development of combative online discourses such as toxicity and political polarization, observable on many social media platforms, such as X (formerly known as Twitter). While this novel pragmatics of online discourse has been widely noted and studied from both the computational and the discourse analysis standpoints, the ideological contents of the discourse have been less scrutinized. In this paper we present an ensemble deep learning classification pipeline able to categorize the political or non-political nature of a short text, as well as the left-right ideological orientation of the text on two separate granularity scales–3-point and 7-point. These models were trained on an existing political ideology dataset which we have compiled from publicly available news websites exhibiting a certain political affiliation. Using BERT-based classifiers, an F1 score of 96.39% is obtained for the political classification subtask, whereas for the coarse-grained and fine-grained ideology classification subtasks, F1 scores of 92.84% and 90.33% are obtained, respectively, on the validation dataset with 5-fold cross-validation. The model was then used to predict the political or non-political nature and the ideological orientation of a sample of 2.6M tweets retrieved from the #Election2020 dataset, which was followed up with an n-gram analysis of the classification results, using s-BERT embeddings, PCA dimensionality reduction and DBSCAN clustering to reveal discussion topics relevant to different ideological orientations. We analyze the results and show the prevalence of each ideology in the data and use SHAP explainers to verify the accuracy of the predictions, revealing stylistic features characteristic of each political orientation that can prove helpful in distinguishing politically-charged conversations, especially extremist discourses.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kovacs et al. (2026) studied this question.

synapsesocial.com/papers/69f5951171405d493affffaehttps://doi.org/10.1109/access.2026.3684000
Ask AI
Helpful
Bookmark
Share
View Full Paper