PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 29, 20250 citationsOpen Access

Alignment of large language models with constrained learning

View Full Paper
BZBotong ZhangSLShuo LiDe Montfort UniversityIHIgnacio HounieCalifornia University of Pennsylvania

Key Points

  • An iterative dual-based alignment method optimizes large language model policies while satisfying constraints.
  • The dual-based approach reduces the optimality gap in the large language model parameter space.
  • Through extensive experiments on the PKU-SafeRLHF dataset, the method demonstrates significant improvements in constrained learning.
  • The study shows that Lagrangian duality enhances convergence and optimality in large language model policy searches.

Abstract

We study the problem of computing an optimal large language model (LLM) policy for a constrained alignment problem, where the goal is to maximize a primary reward objective while satisfying constraints on secondary utilities. Despite the popularity of Lagrangian-based LLM policy search in constrained alignment, iterative primal-dual methods often fail to converge, and non-iterative dual-based methods do not achieve optimality in the LLM parameter space. To address these challenges, we employ Lagrangian duality to develop an iterative dual-based alignment method that alternates between updating the LLM policy via Lagrangian maximization and updating the dual variable via dual descent. In theory, we characterize the primal-dual gap between the primal value in the distribution space and the dual value in the LLM parameter space. We further quantify the optimality gap of the learned LLM policies at near-optimal dual variables with respect to both the objective and the constraint functions. These results prove that dual-based alignment methods can find an optimal constrained LLM policy, up to an LLM parametrization gap. We demonstrate the effectiveness and merits of our approach through extensive experiments conducted on the PKU-SafeRLHF dataset.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68da5a3ec1728099cfd119cehttps://doi.org/10.48550/arxiv.2505.19387
Ask AI
Helpful
Bookmark
Share
View Full Paper