PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 4, 20241 citationsOpen Access

Adaptive Preference Scaling for Reinforcement Learning with Human Feedback

View Full Paper
IHIlgee HongZLZichong LiABAlexander Bukharin

Key Points

  • The proposed adaptive preference loss leads to a more flexible reward function in reinforcement learning, improving policy performance.
  • Analyses revealed that smaller scaling parameters are assigned to ambiguous preferences while larger parameters are used for clear preferences.
  • Methodology employs distributionally robust optimization to optimize the adaptive scaling parameter for each pair of preferences efficiently, ensuring convexity in the loss function approach. The findings indicate that optimizing reward functions can significantly align better with policy optimization and facilitate hyperparameter tuning.

Abstract

Reinforcement learning from human feedback (RLHF) is a prevalent approach to align AI systems with human values by learning rewards from human preference data. Due to various reasons, however, such data typically takes the form of rankings over pairs of trajectory segments, which fails to capture the varying strengths of preferences across different pairs. In this paper, we propose a novel adaptive preference loss, underpinned by distributionally robust optimization (DRO), designed to address this uncertainty in preference strength. By incorporating an adaptive scaling parameter into the loss for each pair, our method increases the flexibility of the reward function. Specifically, it assigns small scaling parameters to pairs with ambiguous preferences, leading to more comparable rewards, and large scaling parameters to those with clear preferences for more distinct rewards. Computationally, our proposed loss function is strictly convex and univariate with respect to each scaling parameter, enabling its efficient optimization through a simple second-order algorithm. Our method is versatile and can be readily adapted to various preference optimization frameworks, including direct preference optimization (DPO). Our experiments with robotic control and natural language generation with large language models (LLMs) show that our method not only improves policy performance but also aligns reward function selection more closely with policy optimization, simplifying the hyperparameter tuning process.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hong et al. (2024) studied this question.

synapsesocial.com/papers/68e664a3b6db6435875f0bbehttps://doi.org/10.48550/arxiv.2406.02764
Ask AI
Helpful
Bookmark
Share
View Full Paper