The autonomous management of Earth Observation (EO) satellites in Low Earth Orbit (LEO) presents a multi-objective optimization challenge under strict operational, energetic, and environmental constraints. This paper introduces a novel, high-fidelity closed-loop simulation framework designed to train and validate a Double Deep Q-Network (DDQN) agent for the autonomous operation of a 6U CubeSat. The proposed environment departs from conventional grid-world abstractions by incorporating an aerospace-grade physics engine built upon a J2-perturbed Keplerian propagator with atmospheric drag, inspired by the SGP4 formulation augmented with J2 short-period perturbations, atmospheric drag modeling, and exact solar ephemeris computation. Navigational realism is preserved through an Extended Kalman Filter (EKF) employing a Runge-Kutta 4 (RK4) integrator and an analytical gravitational Jacobian, processing noisy GPS measurements at a 20-meter standard deviation to produce the estimated orbital state. The satellite middleware explicitly models Newtonian attitude dynamics with Reaction Wheel (RW) saturation and Magnetorquer (MTQ) desaturation against Earth's computed magnetic field. The power subsystem employs a non-linear Shepherd battery model coupled with Wöhler chemical degradation, while a three-node thermodynamic network governs payload thermal throttling. The DDQN agent operates within a 19-dimensional continuous observation space and selects from four discrete operational modes: Sun Pointing, Imaging, Downlink, and Edge AI Processing. A composite reward function balances data monetization from dynamically generated ‘VIP’ targets (high-value strategic maritime chokepoints), payload ‘freshness’ decay, and penalizations for hardware stress. Training proceeds in two phases: an initial global exploration run followed by a Curriculum Learning fine-tuning phase with compressed exploration. The final policy demonstrates robust convergence to a stable, profit-maximizing operational regime, successfully transitioning from stochastic exploration to a consistent exploitation strategy that prioritizes mission longevity and data value. Long-term certification over 10, 000 simulated orbits confirms zero hardware emergencies for the ASRM agent. While the mission naturally concludes at orbit 5, 569 due to atmospheric drag, the agent generates a cumulative revenue of 14. 2 million. In contrast, traditional deterministic industrial heuristics fail to achieve commercial break-even, yielding negligible revenue due to an inability to navigate atmospheric interference and ground station operating costs. The agent demonstrates advanced thermal duty-cycling, strategically leveraging Edge AI processing (14. 4% duty cycle) to achieve a 4: 1 data compression ratio. These results prove that autonomous DRL architectures are not merely an optimization, but a fundamental requirement for the commercial viability of complex, long-term space operations in highly constrained environments.
Matteo Cantelli (Mon,) studied this question.