PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 31, 2026Software Practice and Experience0 citations

QRAP: A Quantum Resource Allocation Platform for Adaptive Scheduling Under Topology Constraints

View Full Paper
SGSong GuoZLZihan LiuZQZhongle Qu

Key Points

  • The aim is to improve task scheduling and resource allocation in noisy intermediate-scale quantum (NISQ) systems.
  • Proposed a reinforcement learning framework for scheduling in distributed NISQ systems.
  • Formulated the scheduling as a constrained optimization problem within a Markov decision process.
  • Implemented deep Q-network and proximal policy optimization agents, comparing against heuristic and random methods.
  • Proximal policy optimization consistently outperformed DQN and heuristic methods.
  • Achieved higher task completion rates with fewer deadline violations.
  • Demonstrated robust adaptation across different reward configurations.

Abstract

ABSTRACT Quantum computing has entered the noisy intermediate‐scale quantum (NISQ) era, where limited qubit numbers, short coherence times, and high error rates pose significant challenges to reliable large‐scale execution. Efficient scheduling and resource allocation across heterogeneous quantum hardware are therefore crucial for maximizing system throughput, fidelity, and fairness. In this work, we propose a hardware‐aware reinforcement learning framework for quantum task scheduling in distributed NISQ systems. Our design explicitly models qubit‐level variability, including connectivity degree, coherence times, error rates, and throughput, while integrating task‐level constraints such as deadlines, priorities, and concurrency requirements. We formulate the scheduling problem as a constrained optimization task and instantiate it as a Markov decision process (MDP), enabling reinforcement learning agents to learn adaptive strategies. Specifically, we implement deep Q‐network (DQN) and proximal policy optimization (PPO) agents, and compare them against heuristic and random baselines. Experimental results demonstrate that PPO consistently outperforms DQN and heuristic methods, achieving higher task completion rates, fewer deadline violations, and more robust adaptation across different reward configurations. This work bridges quantum hardware modeling with reinforcement learning‐based scheduling, providing a practical pathway for resource optimization in distributed quantum computing environments.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Guo et al. (2026) studied this question.

synapsesocial.com/papers/6a1bd2845783ba022b6fdfc1https://doi.org/10.1002/spe.70080
Ask AI
Helpful
Bookmark
Share
View Full Paper