PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 14, 2026International Journal of Advanced Robotic Systems0 citationsOpen Access

Factorial designed experiment to select reinforcement learning observations for a hexapod robot trajectory following task

View Full Paper
AFAlec FreemanRBRobert Bauer

Key Points

  • This research aims to identify the best observations for a reinforcement learning agent to improve a hexapod robot's trajectory-following performance.
  • Developed a hexapod robot simulator in MATLAB Simscape.
  • Utilized a central pattern generator with coupled Hopf oscillators for movement.
  • Applied deep deterministic policy gradient for training the reinforcement learning agent.
  • Conducted quarter-fraction and eighth-fraction factorial designed experiments to test observation combinations.
  • Formulated regression models to predict optimal observation selections.
  • Identified maximum rewarding observations: joint torques, body linear velocities, body orientation, body angular velocities, and body height.
  • Reduced the number of experimental runs from 128 to more manageable combinations.

Abstract

This article explores the use of fractional factorial designed experiments to help select the observations that are provided to a reinforcement learning agent for a hexapod robot trajectory-following task. A hexapod robot simulator is developed in the MATLAB Simscape environment and uses a central pattern generator consisting of 6 coupled Hopf oscillators and corresponding joint angle mapping functions to move the robot. The reinforcement learning agent is trained to control the hexapod using the deep deterministic policy gradient algorithm on a trajectory-following task. To test different combinations of seven potential observations, both quarter-fraction and eighth-fraction factorial designed experiments are proposed to reduce the number of runs from the maximum possible 128. Through the implementation of these designed experiments, regression models were formulated to predict which combinations of observations maximize the hexapod training reward. Model predictions were then validated using the simulator, and the corresponding trajectory-following capabilities of the hexapod were demonstrated. For the conditions used in this research, the observations that obtained the maximum final average reward are the hexapod's joint torques, body linear velocities, body orientation, body angular velocities, and body height above the ground.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Freeman et al. (2026) studied this question.

synapsesocial.com/papers/69b4fc1fb39f7826a300cd6fhttps://doi.org/10.1177/17298806261431893
Ask AI
Helpful
Bookmark
Share
View Full Paper