PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 2, 20250 citationsOpen Access

Intersection of Reinforcement Learning and Bayesian Optimization for Intelligent Control of Industrial Processes: A Safe MPC-based DPG using Multi-Objective BO

View Full Paper
HEHossein Nejatbakhsh EsfahaniJVJavad Mohammadpour Velni

Key Points

  • Integrating MPC with multi-objective bayesian optimization improves control systems' performance and safety during adaptation.
  • The method demonstrated sample-efficient learning while maintaining stability and high performance in a numerical example.
  • Utilizing the expected hypervolume improvement function helps address challenges faced in traditional MPC-RL approaches.
  • The proposed approach achieves better parameter tuning even in the presence of model imperfections.

Abstract

Model Predictive Control (MPC)-based Reinforcement Learning (RL) offers a structured and interpretable alternative to Deep Neural Network (DNN)-based RL methods, with lower computational complexity and greater transparency. However, standard MPC-RL approaches often suffer from slow convergence, suboptimal policy learning due to limited parameterization, and safety issues during online adaptation. To address these challenges, we propose a novel framework that integrates MPC-RL with Multi-Objective Bayesian Optimization (MOBO). The proposed MPC-RL-MOBO utilizes noisy evaluations of the RL stage cost and its gradient, estimated via a Compatible Deterministic Policy Gradient (CDPG) approach, and incorporates them into a MOBO algorithm using the Expected Hypervolume Improvement (EHVI) acquisition function. This fusion enables efficient and safe tuning of the MPC parameters to achieve improved closed-loop performance, even under model imperfections. A numerical example demonstrates the effectiveness of the proposed approach in achieving sample-efficient, stable, and high-performance learning for control systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Esfahani et al. (2025) studied this question.

synapsesocial.com/papers/68de5da783cbc991d0a20cbbhttps://doi.org/10.48550/arxiv.2507.09864
Ask AI
Helpful
Bookmark
Share
View Full Paper