PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 24, 2026Journal of Computational Design and Engineering0 citationsOpen Access

Air combat joint strategy learning based on a dual-loop framework and hindsight experience replay

View Full Paper
YZYuhe ZhangZYZhen YangBZBao Zhang

Key Points

  • This research aims to improve air combat decision-making methods using artificial intelligence by addressing complexities in action selection and reward function design.
  • Developed a novel algorithm based on a dual-loop framework for decision-making.
  • Separated maneuvering and missile launching decisions into distinct optimization processes.
  • Utilized hindsight experience replay to train missile launching decisions and enhance learning samples.
  • Conducted experiments with a self-play agent and air combat bot in a simulation environment.
  • The proposed method generated a joint strategy for maneuvering and missile launching actions.
  • Achieved a higher win rate in adversarial experiments compared to state-of-the-art methods.
  • Demonstrated effective optimization and decision-making in complex air combat scenarios.

Abstract

Abstract The research on air combat decision-making methods using artificial intelligence has become a widely studied field. However, due to the complexity of the air combat process and the problem of hybrid action selection (discrete/continuous), traditional methods struggle to simultaneously make decisions on continuous maneuvering and discrete missile launching actions. In addition, designing complex dense reward functions requires difficult-to-obtain aviation expert knowledge, while relying on sparse reward functions makes it difficult to fully explore a large state space. In view of this, we propose a novel algorithm based on a dual-loop framework. The core idea is to separate maneuvering and missile launching decisions into two optimization processes within the training loop, enabling joint decision-making during the search phase while allowing independent optimization during the optimization phase. Besides, hindsight experience replay is adopted to train missile launching decisions. It expands valuable learning samples through a sample relabeling approach. We designed a series of experiments to validate the performance of the proposed method by constructing the opponent’s strategy using a self-play agent and an air combat bot. The performance of the proposed method was validated in a simulation environment, demonstrating that it can generate an air combat joint strategy incorporating both maneuvering and missile launching. In adversarial experiments, the air combat joint strategy we generated achieved a higher win rate than other state-of-the-art air combat methods.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/6974610cbb9d90c67120ae39https://doi.org/10.1093/jcde/qwag006
Ask AI
Helpful
Bookmark
Share
View Full Paper