PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 23, 2026Autonomous Intelligent Systems0 citationsOpen Access

UAV swarm communication networking and routing optimization for high-demand users: a graph attention multi-agent reinforcement learning approach

ZNZhaopeng NingGLGang LiWLWei Li

Key Points

  • This research aims to optimize UAV swarm communication for high-demand users by addressing trajectory, topology, and routing challenges.
  • Models the problem as a multi-agent partially observable Markov decision process
  • Proposes a graph attention-based multi-agent deep deterministic policy gradient algorithm
  • Considers user quality of service, system throughput, and energy constraints in the reward function
  • Achieved 50% improvement in convergence speed
  • Increased user service satisfaction by 12% to 18%
  • Improved system throughput by 25% to 40%

Abstract

Unmanned aerial vehicle swarms serving ground high-demand communication users in dynamic environments must simultaneously optimize three-dimensional trajectories, communication network topology, and routing strategies while considering limited energy, link quality fluctuations, and collision avoidance constraints. This problem faces three core challenges: routing decisions under dynamic topology require real-time adaptation to vehicle position changes and channel variations; end-to-end delay and throughput optimization in multi-hop communication demands coordinated forwarding strategies across all vehicles; high-dimensional continuous action spaces and partial observability make traditional optimization methods difficult to solve. This paper models the problem as a multi-agent partially observable Markov decision process and proposes a graph attention-based multi-agent deep deterministic policy gradient algorithm to jointly optimize velocity vectors, communication power, and routing decisions for each vehicle. The reward function comprehensively considers user quality of service, system throughput, end-to-end delay, and energy consumption while ensuring safety distance and energy margins through constraint penalties. Simulation results demonstrate that compared to single-agent deep deterministic policy gradient and independent Q-learning baseline methods, the proposed method achieves approximately 50% improvement in convergence speed, 12% to 18% increase in user service satisfaction, 25% to 40% improvement in system throughput, 30% to 45% reduction in end-to-end delay, and 39% to 102% improvement in energy efficiency. The framework dynamically adjusts network topology and routing strategies according to user demands, providing a deployable solution for large-scale vehicle swarm communication networks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ning et al. (2026) studied this question.

synapsesocial.com/papers/69e9b71b85696592c86eb20fhttps://doi.org/10.1007/s43684-026-00131-6
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Joint Trajectory Control, Frequency Allocation, and Routing for UAV Swarm Networks: A Multi-Agent Deep Reinforcement Learning Approach2024 · 9 citations
  2. 2A Hybrid Communication Method for UAV Swarms Based on ADG-MASAC2026
  3. 3Modeling Dynamic Obstacle Avoidance Strategy of Drone Swarms Combined with Multi-Agent Reinforcement Learning2026
  4. 4Graph Attention-Based Reinforcement Learning for Trajectory Design and Resource Assignment in Multi-UAV Assisted Communication2024 · 44 citations
  5. 5Multi-objective Aerial Collaborative Secure Communication Optimization via Generative Diffusion Model-enabled Deep Reinforcement Learning2024