PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

Toward Preference-aligned Large Language Models via Residual-based Model Steering

View Full Paper
LCLucio La CavaATAndrea Tagarelli

Key Points

  • PaLRS offers significant improvements in preference alignment for large language models, enhancing their usefulness without extensive training.
  • Models using PaLRS show better performance in mathematical reasoning and code generation benchmarks compared to traditional methods.
  • Evaluation across various small-to-medium-scale open-source LLMs highlights time efficiency and flexible integration of the method.
  • PaLRS reduces reliance on curated data, demonstrating efficiency while maintaining general-purpose performance in LLMs.

Abstract

Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences. Existing approaches such as Reinforcement Learning from Human Feedback or Direct Preference Optimization typically require curated data and expensive optimization over billions of parameters, and eventually lead to persistent task-specific models. In this work, we introduce Preference alignment of Large Language Models via Residual Steering (PaLRS), a training-free method that exploits preference signals encoded in the residual streams of LLMs. From as few as one hundred preference pairs, PaLRS extracts lightweight, plug-and-play steering vectors that can be applied at inference time to push models toward preferred behaviors. We evaluate PaLRS on various small-to-medium-scale open-source LLMs, showing that PaLRS-aligned models achieve consistent gains on mathematical reasoning and code generation benchmarks while preserving baseline general-purpose performance. Moreover, when compared to DPO-aligned models, they perform better with huge time savings. Our findings highlight that PaLRS offers an effective, much more efficient and flexible alternative to standard preference optimization pipelines, offering a training-free, plug-and-play mechanism for alignment with minimal data.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cava et al. (2025) studied this question.

synapsesocial.com/papers/68f5fcce8d54a28a75cf1ae2https://doi.org/10.48550/arxiv.2509.23982
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization2024 · 1 citations
  2. 2Active Preference Learning for Large Language Models2024
  3. 3Aligning Large Language Models with Self-generated Preference Data2024
  4. 4Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization2025
  5. 5Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators2024 · 10 citations