PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 10, 2026Aerospace Science and Technology0 citationsOpen Access

Meta Reinforcement Learning Method for Dynamic Mission Scheduling of Earth Observation Satellites

View Full Paper
WYWei YaoXSXin ShenGZGuo Zhang

Key Points

  • The aim is to develop a meta reinforcement learning framework for dynamic mission scheduling of Earth observation satellites.
  • Implemented a Markov process for dynamic mission environments
  • Integrated adaptive reward mechanisms and dynamic action spaces
  • Used long short-term memory networks for policy generalization
  • Applied proximal policy optimization algorithm as the inner loop
  • Evaluated on satellite orbit and hybrid mission datasets
  • MLR-DMS outperformed traditional methods like PPO, A2C, and DQN
  • Significant increases in mission completion rates and cumulative rewards
  • Enhanced computational efficiency under various mission densities
  • Robust adaptability in dynamic and uncertain operational environments

Abstract

• Markov process design adapted for dynamic satellite mission environments • Adaptive reward mechanism for dynamic mission scheduling • Meta-RL for hybrid mission generalization Earth observation satellites (EOSs) play crucial roles in disaster monitoring, resource management, military reconnaissance, and environmental protection. However, the increasing complexity and dynamic nature of EOS missions pose significant challenges for conventional mission scheduling methods, which frequently struggle with high computational overhead, poor adaptability to real-time changes, and limited generalizability across mission scenarios. Thus, this paper proposed a meta reinforcement learning (MRL) framework for hybrid dynamic mission scheduling (DMS) for EOSs. The proposed MRL-DMS method integrates a mission-adaptive reward mechanism and a dynamic action space within a Markov decision process formulation to enable real-time responsiveness to stochastic mission arrivals. A metalearning layer based on long short-term memory networks is implemented to enhance policy generalization across diverse mission distributions. Leveraging proximal policy optimization (PPO) as the inner loop reinforcement learning algorithm, the MRL-DMS method adapts to new mission environments with minimal retraining. The proposed method was evaluated on realistic satellite orbit data and largescale hybrid mission datasets derived from global conflict scenarios. The experimental results demonstrate that the MRL-DMS method outperforms the state-of-the-art PPO, A2C, and DQN algorithms. The MRL-DMS method achieves significant improvements in the dynamic mission completion rates, cumulative reward acquisition, and computational efficiency. In addition, MRL-DMS exhibits robust adaptability across varying mission densities, scales, and spatial distributions, effectively prioritizing high-urgency missions while maintaining stable performance under scheduling pressure. The MRL-DMS method provides a scalable, intelligent solution for autonomous satellite mission planning in dynamic and uncertain operational environments. The findings of this study indicate that the MRL-DMS method provides valuable insights into real-time remote sensing strategies in response to global emergencies and rapidly evolving geopolitical events.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yao et al. (2026) studied this question.

synapsesocial.com/papers/69af95b470916d39fea4d98bhttps://doi.org/10.1016/j.ast.2026.112094
Ask AI
Helpful
Bookmark
Share
View Full Paper