PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 10, 2026Discover Applied Sciences0 citationsOpen Access

Quantifying modality imbalance and visual jailbreak robustness in LLaVA via projected gradient descent

SASaklain AbdullahRHRiad HossainMCMahfuzulhoq Chowdhury

Key Points

  • This research aims to assess the vulnerability of LLaVA-1.5 to adversarial attacks using targeted visual prompts.
  • Applied Projected Gradient Descent (PGD) attack on 1000 samples from the MM-SafetyBench dataset.
  • Evaluated high risk categories to quantify attack effectiveness.
  • Implemented strict metrics for compliance to eliminate false positives.
  • Achieved Attack Success Rates (ASR) of 95% to 100% at perturbation budgets of ε ≥ 8/255 across all categories.
  • Adversarial visual embeddings successfully overpowered textual safety constraints.
  • Findings indicate a significant modality gap where visual inputs can subvert safety mechanisms.

Abstract

While Large Vision Language Models (LVLMs) exhibit remarkable capabilities, their visual modality introduces a critical attack surface that can bypass text only safety alignments. This paper evaluates the vulnerability of LLaVA-1. 5 to targeted adversarial visual prompts designed to induce malicious compliance. Using a Projected Gradient Descent (PGD) attack on the MM-SafetyBench dataset, we evaluate 1000 samples across five high risk categories. To eliminate false positives caused by superficial compliance, we apply a rigorous metric that strictly demands sustained, direct compliance without late stage refusals. Our results demonstrate that imperceptible visual perturbations effectively hijack safety guardrails, achieving Attack Success Rates (ASR) of 95% to 100% across all categories at perturbation budgets of 8/255. Furthermore, analysis of the Modality Gap (₌₆) reveals that adversarial visual embeddings overpower textual safety constraints, forcing a malicious multimodal alignment. These findings underscore the inadequacy of current unimodal safety fine tuning and highlight the urgent need for robust, multimodal specific defense mechanisms.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Abdullah et al. (2026) studied this question.

synapsesocial.com/papers/6a002147c8f74e3340f9c1c6https://doi.org/10.1007/s42452-026-08793-w
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Real-Time Deepfake Video Detection Using Eye Movement Analysis with a Hybrid Deep Learning Approach2024 · 69 citations
  2. 2Visual Instruction Tuning2023 · 904 citations
  3. 3Training Language Models to Follow Instructions with Human Feedback2022 · 981 citations
  4. 4FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts2025 · 53 citations
  5. 5Enhancing multimodal deepfake detection with local–global feature integration and diffusion models2025 · 46 citations