PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 7, 2024Robotic Intelligence and Automation1 citations

Multi-objective reinforcement learning based on nonlinear scalarization and long-short-term optimization

View Full Paper
HWHongze Wang

Key Points

Key points are not available for this paper at this time.

Abstract

Purpose Many practical control problems require achieving multiple objectives, and these objectives often conflict with each other. The existing multi-objective evolutionary reinforcement learning algorithms cannot achieve good search results when solving such problems. It is necessary to design a new multi-objective evolutionary reinforcement learning algorithm with a stronger searchability. Design/methodology/approach The multi-objective reinforcement learning algorithm proposed in this paper is based on the evolutionary computation framework. In each generation, this study uses the long-short-term selection method to select parent policies. The long-term selection is based on the improvement of policy along the predefined optimization direction in the previous generation. The short-term selection uses a prediction model to predict the optimization direction that may have the greatest improvement on overall population performance. In the evolutionary stage, the penalty-based nonlinear scalarization method is used to scalarize the multi-dimensional advantage functions, and the nonlinear multi-objective policy gradient is designed to optimize the parent policies along the predefined directions. Findings The penalty-based nonlinear scalarization method can force policies to improve along the predefined optimization directions. The long-short-term optimization method can alleviate the exploration-exploitation problem, enabling the algorithm to explore unknown regions while ensuring that potential policies are fully optimized. The combination of these designs can effectively improve the performance of the final population. Originality/value A multi-objective evolutionary reinforcement learning algorithm with stronger searchability has been proposed. This algorithm can find a Pareto policy set with better convergence, diversity and density.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hongze Wang (2024) studied this question.

synapsesocial.com/papers/68e6b3a7b6db6435876348a3https://doi.org/10.1108/ria-11-2023-0174
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1MORSE: Multi-Objective Reinforcement Learning via Strategy Evolution for Supply Chain Optimization2025
  2. 2NeuroAction: a neuroevolutionary approach to reinforcement learning for autonomous vehicles2026
  3. 3Discovering a Single Neural Network Controller for Multiple Tasks with Evolutionary Algorithms2025
  4. 4A many-objective evolutionary algorithm based on learning assessment and mapping guidance of historical superior information2024 · 7 citations
  5. 5Large-Scale Sparse Multimodal Multiobjective Optimization via Multi-Stage Search and RL-Assisted Environmental Selection2026 · 1 citations