In this note, we present the first analysis of the computational speedup achieved using the GPU version of Madgraph, known as Madgraph4GPU (MG4GPU), in the CMS workflow. Madgraph is one of the most widely used event generators in CMS. This work represents the initial step toward benchmarking the improvements offered by both the GPU and vectorized CPU implementations. We demonstrate timing improvements across a broad range of physics processes relevant to CMS. Speedups are quantified for both gridpack production and event generation. A gridpack is a pre-defined package that encapsulates all the necessary components for effectively executing Monte Carlo event simulations, eliminating redundant computations of common elements for each event. Preliminary results indicate a speedup of approximately a factor of three with vectorized CPUs and an order-of-magnitude improvement with GPUs in gridpack production for the Drell-Yan and top quark pair production processes. For event generation, we observe a speedup of 1.5 times with vectorized CPUs and 7 times with GPUs when generating 105 events. These workflows were tested using a variety of computational resources, including CUDA-enabled NVIDIA GPUs and modern vectorized CPUs from Intel and AMD, accessible via CERN resources and HPCs.
Choi et al. (Tue,) studied this question.