PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20250 citationsOpen Access

Learning the Topic, Not the Language: How LLMs Classify Online Immigration Discourse Across Languages

View Full Paper
ANAndrea NasutoSIStefano M. IacusFRFrancisco Rowe

Key Points

  • LLMs fine-tuned in one or two languages can classify immigration-related content across unseen languages.
  • Minimal language-specific fine-tuning shows significant gains in topic detection with under-represented languages.
  • Multilingual fine-tuning enhances identification of pro- or anti-immigration sentiment in tweets.
  • The study presents lightweight interventions to correct pre-training bias favoring dominant languages.

Abstract

Large language models (LLMs) are transforming social-science research by enabling scalable, precise analysis. Their adaptability raises the question of whether knowledge acquired through fine-tuning in a few languages can transfer to unseen languages that only appeared during pre-training. To examine this, we fine-tune lightweight LLaMA 3. 2-3B models on monolingual, bilingual, or multilingual data sets to classify immigration-related tweets from X/Twitter across 13 languages, a domain characterised by polarised, culturally specific discourse. We evaluate whether minimal language-specific fine-tuning enables cross-lingual topic detection and whether adding targeted languages corrects pre-training biases. Results show that LLMs fine-tuned in one or two languages can reliably classify immigration-related content in unseen languages. However, identifying whether a tweet expresses a pro- or anti-immigration stance benefits from multilingual fine-tuning. Pre-training bias favours dominant languages, but even minimal exposure to under-represented languages during fine-tuning (as little as 9. 6210^-11 of the original pre-training token volume) yields significant gains. These findings challenge the assumption that cross-lingual mastery requires extensive multilingual training: limited language coverage suffices for topic-level generalisation, and structural biases can be corrected with lightweight interventions. By releasing 4-bit-quantised, LoRA fine-tuned models, we provide an open-source, reproducible alternative to proprietary LLMs that delivers 35 times faster inference at just 0. 00000989% of the dollar cost of the OpenAI GPT-4o model, enabling scalable, inclusive research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nasuto et al. (2025) studied this question.

synapsesocial.com/papers/68f12bfb2107091eab27a367https://doi.org/10.48550/arxiv.2508.06435
Ask AI
Helpful
Bookmark
Share
View Full Paper