PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 30, 20260 citationsOpen Access

Multiclass abusive language detection task – challenging Transformers with RNNs models

View Full Paper
SRStefano RinaudoBVBeatrice Vandi

Key Points

  • The study explores if RNN models can achieve comparability in performance to transformers for multiclass abusive language detection.
  • Utilized AbuseEval v1.0 and OLID datasets for model training.
  • Implemented preprocessing techniques such as tokenization and punctuation filtering.
  • Tested data augmentation and various dropout methods to address class imbalance.
  • Evaluated multiple LSTM and BiLSTM configurations, selecting the best-performing model based on macro F1 score.
  • Model 4, a BiLSTM, achieved a macro F1 score of approximately 0.51, close to the BERT baseline of 0.535.
  • Both RNN and transformer architectures were effective in identifying explicit abusive and non-abusive tweets.
  • Implicit abusive language detection remains challenging, suggesting room for future research.

Abstract

In this report, we investigated whether a carefully tuned RNN-based model can match Transformer performance on multiclass abusive language detection in tweets. Using the AbuseEval v1. 0 / OLID datasets, we applied consistent preprocessing (tokenization, punctuation filtering, placeholder removal), vocabulary-limited vectorization, and used GloVe-Twitter embeddings. To mitigate class imbalance and improve minority-class learning we tested data augmentation (using contextual BERT model), focal loss, recurrent and spatial dropout, and class weighting. We experimented with several LSTM and BiLSTM configurations and selected Model 4, a BiLSTM with SpatialDropout1D (0. 3), recurrent dropout (0. 5) and focal loss (γ=2. 0), as the best trade-off between stability and class-wise performance. Model 4 achieved a macro F1 ≈ 0. 51, comparable to the BERT baseline reported by Caselli et al. (2020): macro F1 ≈ 0. 535. Both architectures perform well on recognizing not abusive and explicitly abusive tweets, but implicit abuse detection remains problematic. Results suggest that an optimized RNN can be competitive with an older BERT baseline while being computationally lighter; future work should evaluate modern Transformer variants and richer contextual methods to better capture implicit abuse. References Caselli T. , Basile V. , Mitrović J. , Kartoziya I. , Granitzer M. (2020). I Feel Offended, Don’t Be Abusive! Implicit/Explicit Messages in Offensive and Abusive Language. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pp. 6193–6202, Marseille, France. European Language Resources Association. Link: https: //aclanthology. org/2020. lrec-1. 760/ Devlin, J. , Chang, M. -W. , Lee, K. , Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, June. Association for Computational Linguistics. DOI: 10. 18653/v1/N19-1423 ElSherief M. , Ziems C. , Muchlinski D. , Anupindi V. , Seybolt J. , De Choudhury M. , Yang D. (2021). Latent hatred: A benchmark for understanding implicit hate speech. arXiv. DOI: 10. 48550/arXiv. 2109. 05322 Hafeez F. , Hafeez M. , Shaff Bin Imran M. , Qasim A. , Hussain N. , Ahmad F. , Sidorov G. (2026). A Comparative Study of Deep Learning and Transformer Models for Twitter Sentiment Analysis. In International Journal of Combinatorial Optimization Problems and Informatics, 17 (1), pp. 85–89. DOI: 10. 61467/2007. 1558. 2026. v17i1. 1241 Mandal R. , Chen J. , Becken S. , Stantic B. (2021). Empirical Study of Tweets Topic Classification Using Transformer-Based Language Models. In Nguyen, N. T. , Chittayasothorn, S. , Niyato, D. , Trawiński, B. (eds) Intelligent Information and Database Systems. ACIIDS 2021. Lecture Notes in Computer Science, vol 12672. Springer, Cham. DOI: 10. 1007/978-3-030-73280-6₂7 Pennington J. , Socher R. , Manning C. D. (2014). GloVe: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1532–1543, Doha, Qatar. Association for Computational Linguistics. DOI: 10. 3115/v1/D14-1162 Trivedi S. (2020), Understanding Focal Loss – A Quick Read. In VisionWizard. Link: https: //medium. com/visionwizard/understanding-focal-loss-a-quick-read-b914422913e7 Vaswani A. , Shazeer N. , Parmar N. , Uszkoreit J. , Jones L. , Gomez A. N. , Kaiser Ł. , Polosukhin I. (2017). Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, & R. Garnett (Eds. ), Advances in neural information processing systems (Vol. 30). Curran Associates, Inc. DOI: 10. 48550/arXiv. 1706. 03762 Zampieri M. , Malmasi S. , Nakov P. , Rosenthal S. , Farra N. , Kumar, R. (2019). SemEval-2019 Task 6: Identifying and categorizing offensive language in social media (OffensEval). In Proceedings of the 13th International Workshop on Semantic Evaluation. pp. 75–86, Minneapolis, Minnesota, USA. Association for Computational Linguistics. DOI: 10. 18653/v1/S19-2010

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rinaudo et al. (2026) studied this question.

synapsesocial.com/papers/69f2a4b78c0f03fd67763d23https://doi.org/10.5281/zenodo.19847349
Ask AI
Helpful
Bookmark
Share
View Full Paper