PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 30, 20250 citationsOpen Access

Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO

View Full Paper
RSRuizhe ShiMSMin‐Kyu SongRZRunlong Zhou

Key Points

  • The performance gap between RLHF and DPO arises from explicit and implicit representation gaps under various optimization settings.
  • Exact optimization reveals that the capabilities of reward and policy model classes significantly influence the final policy qualities.
  • In scenarios with isomorphic and mis-specified models, online DPO can outperform both RLHF and standard DPO methods.
  • RLHF requires fewer samples than DPO to develop effective reward models, offering a statistical advantage in learning.

Abstract

We present a fine-grained theoretical analysis of the performance gap between reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO) under a representation gap. Our study decomposes this gap into two sources: an explicit representation gap under exact optimization and an implicit representation gap under finite samples. In the exact optimization setting, we characterize how the relative capacities of the reward and policy model classes influence the final policy qualities. We show that RLHF, DPO, or online DPO can outperform one another depending on the type of model mis-specifications. Notably, online DPO can outperform both RLHF and standard DPO when the reward and policy model classes are isomorphic and both mis-specified. In the approximate optimization setting, we provide a concrete construction where the ground-truth reward is implicitly sparse and show that RLHF requires significantly fewer samples than DPO to recover an effective reward model -- highlighting a statistical advantage of two-stage learning. Together, these results provide a comprehensive understanding of the performance gap between RLHF and DPO under various settings, and offer practical insights into when each method is preferred.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shi et al. (2025) studied this question.

synapsesocial.com/papers/68dc12c58a7d58c25ebb089fhttps://doi.org/10.48550/arxiv.2505.19770
Ask AI
Helpful
Bookmark
Share
View Full Paper