PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 4, 2026Electronics0 citationsOpen Access

Defamiliarization Attack: Literary Theory Enabled Discussion of LLM Safety

View Full Paper
BBBibin BabuIAIana AgafonovaSBSebastian Biedermann

Key Points

  • The aim is to explore how a defamiliarization attack can expose vulnerabilities in large language models (LLMs).
  • Introduced a jailbreaking attack called Defamiliarization with multi-turn queries.
  • Documented scenarios where prompts led to unethical outputs and overlooked critical events.
  • Analyzed relationships between model scales and susceptibility to manipulation.
  • Smaller-parameter models were easier to manipulate using defamiliarized prompts.
  • Demonstrated that existing alignment strategies based only on trigger detection are inadequate.
  • Advocated for a holistic approach to LLM safety that incorporates literary theory and ethics.

Abstract

This paper introduces a multi-turn large language model (LLM) jailbreaking attack called Defamiliarization, in which malicious queries are embedded within ostensibly harmless narratives. By reframing requests in “unmarked” contexts, LLMs can be coerced into producing undesirable outputs. A range of scenarios is documented, from planning ethically dubious actions to selectively overlooking critical events in literary texts, thereby exposing the limitations of alignment strategies predicated on detecting trigger words or semantic cues. Rather than substituting vocabulary, defamiliarization manipulates context and presentation, highlighting vulnerabilities that cannot be addressed by token-level fixes alone. Beyond demonstrating the effectiveness of defamiliarization as an attack strategy, evidence is presented of a systematic relationship between model scale and susceptibility. Experiments reveal that smaller-parameter models are significantly easier to manipulate using defamiliarized prompts. This finding raises important concerns regarding the growing popularity of lightweight, locally hosted LLMs, which are favored for their lower computational requirements but may lack alignment safeguards. A more holistic approach to LLM safety is advocated—one that incorporates insights from literary theory, ethics, and user experience—treating these models as interpretive agents. By doing so, defenses against covert manipulations can be strengthened and AI systems can remain aligned with human values.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Babu et al. (2026) studied this question.

synapsesocial.com/papers/69a7cd3dd48f933b5eed9562https://doi.org/10.3390/electronics15051047
Ask AI
Helpful
Bookmark
Share
View Full Paper