We introduce just-in-time (JIT) compilation to the integral kernels for Gaussian-type orbitals to enhance the efficiency of electron repulsion integral computations. For Coulomb and exchange (JK) matrices, JIT-based algorithms yield a 2× speedup for the small 6-31G* basis set over GPU4PySCF v1.4 on an NVIDIA A100-80G GPU. By incorporating a novel algorithm designed for orbitals with high angular momentum, the efficiency of JK evaluations with the large def2-TZVPP basis set is improved by up to 4×. The core CUDA implementation is compact, comprising only ∼1000 lines of code, including support for single-precision arithmetic. Furthermore, the single-precision implementation achieves a 3× speedup over the previous state-of-the-art.
Wu et al. (2026) studied this question.