This article explores the use of fractional factorial designed experiments to help select the observations that are provided to a reinforcement learning agent for a hexapod robot trajectory-following task. A hexapod robot simulator is developed in the MATLAB Simscape environment and uses a central pattern generator consisting of 6 coupled Hopf oscillators and corresponding joint angle mapping functions to move the robot. The reinforcement learning agent is trained to control the hexapod using the deep deterministic policy gradient algorithm on a trajectory-following task. To test different combinations of seven potential observations, both quarter-fraction and eighth-fraction factorial designed experiments are proposed to reduce the number of runs from the maximum possible 128. Through the implementation of these designed experiments, regression models were formulated to predict which combinations of observations maximize the hexapod training reward. Model predictions were then validated using the simulator, and the corresponding trajectory-following capabilities of the hexapod were demonstrated. For the conditions used in this research, the observations that obtained the maximum final average reward are the hexapod's joint torques, body linear velocities, body orientation, body angular velocities, and body height above the ground.
Freeman et al. (2026) studied this question.