Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
June 26, 2024Open Access

SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance

View Full Paper
Ask AI
Bookmark
Share

Authors

CHCaishuang HuangWZWanxu ZhaoRZRui Zheng

Discussion

Loading...

Member takes

Overview

Key Points

Key points are not available for this paper at this time.

Cite This Study

Huang et al. (2024) studied this question.

synapsesocial.com/papers/68e633aeb6db6435875c5410https://doi.org/10.48550/arxiv.2406.18118
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding2024 · 3 citations
  2. 2How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States2024
  3. 3Alignment-Enhanced Decoding:Defending via Token-Level Adaptive Refining of Probability Distributions2024
  4. 4ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning2025
  5. 5ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack2025