Key points are not available for this paper at this time.
Large Language Models (LLMs) are typically safety-aligned using high-resource language data, but it remains unclear whether these constraints transfer reliably across distinct linguistic manifolds. This study examines the mathematical foundations of cross-lingual guardrail degradation using Bengali as a low-resource test case. We evaluate Meta-Llama-3-8B, Gemma-2-9B, and Llama-Guard-3 through an automated English-to-Bengali translation pipeline, paired statistical testing, latent-space visualization, and tokenization-based structural analysis. The results show a statistically significant increase in Gemma-2’s Attack Success Rate from 32.0% in English to 41.2% in Bengali (p<0.0001, McNemar’s test), while Llama-Guard-3 fails to detect 39.5% of malicious Bengali prompts. Latent-space projections indicate weaker separation between safe and unsafe Bengali representations, and tokenization analysis shows a 4.69-fold token fertility expansion associated with a normalized perplexity of 887.32. Furthermore, projecting low-resource inputs back into the high-resource latent space successfully restores optimization constraints, whereas natively translating safety prompts exacerbates vulnerability. Together, these findings suggest that cross-lingual safety failures are associated with representational entanglement and token fragmentation rather than only superficial prompt translation effects. The study supports the need for multilingual alignment methods that better account for tokenization geometry, latent-space structure, and language-dependent safety evaluation.
Hasan et al. (Tue,) studied this question.