PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 26, 20241 citationsOpen Access

Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models

View Full Paper
TTTianyi TangWLWenyang LuoHHHaoyang Huang

Key Points

Key points are not available for this paper at this time.

Abstract

Large language models (LLMs) demonstrate remarkable multilingual capabilities without being pre-trained on specially curated multilingual parallel corpora. It remains a challenging problem to explain the underlying mechanisms by which LLMs process multilingual texts. In this paper, we delve into the composition of Transformer architectures in LLMs to pinpoint language-specific regions. Specially, we propose a novel detection method, language activation probability entropy (LAPE), to identify language-specific neurons within LLMs. Based on LAPE, we conduct comprehensive experiments on two representative LLMs, namely LLaMA-2 and BLOOM. Our findings indicate that LLMs' proficiency in processing a particular language is predominantly due to a small subset of neurons, primarily situated in the models' top and bottom layers. Furthermore, we showcase the feasibility to "steer" the output language of LLMs by selectively activating or deactivating language-specific neurons. Our research provides important evidence to the understanding and exploration of the multilingual capabilities of LLMs.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tang et al. (2024) studied this question.

synapsesocial.com/papers/68e779e4b6db6435876ee86chttps://doi.org/10.48550/arxiv.2402.16438
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Unraveling Babel: Exploring Multilingual Activation Patterns within Large Language Models2024
  2. 2How do Large Language Models Handle Multilingualism?2024 · 5 citations
  3. 3The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs2025
  4. 4Sharing Matters: Analysing Neurons Across Languages and Tasks in LLMs2024
  5. 5Understanding the role of FFNs in driving multilingual behaviour in LLMs2024