Task-oriented dialogue systems face a tension between comprehensive constraint elicitation (task adequacy) and conversational efficiency (minimizing turns). Current preference learning frameworks treat preferences as static, unable to capture the dynamic evolution of interaction states that evolve across dialogue progression. We present Dual-DPO, a framework that embeds multi-objective preferences into data construction via turn-aware scoring. Our approach decouples objective balancing from policy updates through offline preference scalarization, addressing the optimization instability challenges in online multi-objective reinforcement learning. Experiments on MultiWOZ 2.4 demonstrate 28–35% dialogue turn reduction while maintaining Joint Goal Accuracy > 89% (p<0.001). Pareto frontier analysis shows 94% coverage with hypervolume HV=0.847. Independent expert evaluation by 10 PhD-level researchers (n=300 assessments, inter-rater agreement α=0.78) confirms 32% user satisfaction improvement (p<0.001). Theoretical analysis demonstrates that offline scalarization, which correlates with improved optimization stability, achieves 3.2× lower gradient variance than online multi-reward optimization by eliminating sampling stochasticity. Our approach enables balanced treatment of competing objectives through Pareto-optimal trade-offs. These results highlight a symmetric and balanced treatment of competing objectives within a Pareto-optimal optimization framework.
Bao et al. (Tue,) studied this question.