Bi-level optimization research area has become increasingly popular, largely due to its effectiveness in modeling and solving real-world problems. This framework provides a hierarchical structure involving two decision-makers (i.e., upper and lower levels) that govern together to find an optimal solution to complex optimization problems. Most resolution methods proposed in the literature adhere to this hierarchical structure, which limit their applicability only to small-scale instances of the problem. Among these resolution strategies, we highlight an interesting evolutionary algorithm known as CODBA, which focuses on decomposing the lower-level search space into several parts that evolve in parallel to address the high complexity of the nested structure. In this paper, we enhance the searching capabilities of CODBA by proposing a novel evolutionary reinforcement learning approach that integrates the core CODBA scheme with a Q-learning strategy, presenting a promising method for training intelligent search algorithms for bi-level optimization problems. The computational statistical experiments are performed on bi-level multi-depot vehicle routing problem, demonstrated the effectiveness of our solution approach in terms of computation time and solution quality compared to existing algorithms.
Abir Chaabani (Tue,) studied this question.