JAliEn, the Grid middleware of the ALICE experiment, utilises whole-node scheduling to maximise resource utilisation from participating sites. This approach offers flexibility in resource allocation and partitioning, allowing for customised configurations that adapt to the evolving needs of the experiment. This scheduling model is gaining traction among Grid sites due to its performance benefits. Additionally, understanding common execution patterns for different workloads allows for more efficient scheduling and resource allocation strategies. Managing the entire set of resources on a node requires careful orchestration. To that end, JAliEn employs custom mechanisms to dynamically allocate idle resources to pending workloads, ensuring that overall resource usage stays within the capacity of a node. This paper evaluates the experiences of the first sites using whole-node scheduling. It highlights its suitability for accommodating jobs with varying resource demands, particularly those with high memory requirements.
Ferrer et al. (Tue,) studied this question.