PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 17, 2026Algorithms0 citationsOpen Access

Key-Value Mapping-Based Text-to-Image Diffusion Model Backdoor Attacks

View Full Paper
LCLujia ChaiYHY. Thomas HouGLGuozhao Liao

Key Points

  • The research aims to explore vulnerabilities in text-to-image diffusion models and propose efficient backdoor attack methods.
  • Two backdoor attack methods were developed: AttnBackdoor and SemBackdoor.
  • AttnBackdoor fine-tunes key-value projection matrices in U-Net cross-attention layers.
  • SemBackdoor edits the MLP projection matrix in the text encoder.
  • Both attacks achieve high success rates over 90%, with SemBackdoor at 98.6% and AttnBackdoor at 97.2%.
  • Parameter updates and training time are reduced by 1–2 orders of magnitude compared to existing methods.
  • The attacks highlight vulnerabilities at both visual and semantic levels.

Abstract

Text-to-image (T2I) generation, a core component of generative artificial intelligence(AI), is increasingly important for creative industries and human–computer interaction. Despite impressive progress in realism and diversity, diffusion models still exhibit critical security blind spots particularly in the Transformer key-value mapping mechanism that underpins cross-modal alignment. Existing backdoor attacks often rely on large-scale data poisoning or extensive fine-tuning, leading to low efficiency and limited stealth. To address these challenges, we propose two efficient backdoor attack methods AttnBackdoor and SemBackdoor grounded in the Transformer’s key-value storage principle. AttnBackdoor injects precise mappings between trigger prompts and target instances by fine-tuning the key-value projection matrices in U-Net cross-attention layers (≈5% of parameters). SemBackdoor establishes semantic-level mappings by editing the text encoder’s MLP projection matrix (≈0.3% of parameters). Both approaches achieve high attack success rates (>90%), with SemBackdoor reaching 98.6% and AttnBackdoor 97.2%. They also reduce parameter updates and training time by 1–2 orders of magnitude compared to prior work while preserving benign generation quality. Our findings reveal dual vulnerabilities at visual and semantic levels and provide a foundation for developing next generation defenses for secure generative AI.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chai et al. (2026) studied this question.

synapsesocial.com/papers/696b26d7d2a12237a934a213https://doi.org/10.3390/a19010074
Ask AI
Helpful
Bookmark
Share
View Full Paper