PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 24, 20250 citationsOpen Access

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement

View Full Paper
YSYu ShenZHZisu HuangZGZhengkang Guo

Key Points

  • IntentionReasoner significantly enhances safety in large language models while reducing harmful content generation.
  • Through extensive evaluation, the mechanism improves generation quality and reduces over-refusal rates, achieving notable metrics.
  • The approach employs intent reasoning and query rewriting, demonstrating effectiveness in multiple safeguard benchmarks.
  • Reinforcement learning strategies enhance performance by integrating rule-based heuristics, highlighting a novel technique.

Abstract

The rapid advancement of large language models (LLMs) has driven their adoption across diverse domains, yet their ability to generate harmful content poses significant safety challenges. While extensive research has focused on mitigating harmful outputs, such efforts often come at the cost of excessively rejecting harmless prompts. Striking a balance among safety, over-refusal, and utility remains a critical challenge. In this work, we introduce IntentionReasoner, a novel safeguard mechanism that leverages a dedicated guard model to perform intent reasoning, multi-level safety classification, and query rewriting to neutralize potentially harmful intent in edge-case queries. Specifically, we first construct a comprehensive dataset comprising approximately 163,000 queries, each annotated with intent reasoning, safety labels, and rewritten versions. Supervised fine-tuning is then applied to equip the guard model with foundational capabilities in format adherence, intent analysis, and safe rewriting. Finally, we apply a tailored multi-reward optimization strategy that integrates rule-based heuristics and reward model signals within a reinforcement learning framework to further enhance performance. Extensive experiments show that IntentionReasoner excels in multiple safeguard benchmarks, generation quality evaluations, and jailbreak attack scenarios, significantly enhancing safety while effectively reducing over-refusal rates and improving the quality of responses.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shen et al. (2025) studied this question.

synapsesocial.com/papers/68d6d82e8b2b6861e4c3e091https://doi.org/10.48550/arxiv.2508.20151
Ask AI
Helpful
Bookmark
Share
View Full Paper