With the growing availability of machine-learned interatomic potential (MLIP) models for materials simulations, there is an increasing demand for robust, automated, and chemically informed benchmarking methodologies. In response, we here introduce LiPS-25, a curated benchmark data set for a canonical series of solid-state electrolyte materials from the Li2S-P2S5 pseudobinary compositional line, including crystalline and amorphous configurations. Together with the data set, we present a suite of performance tests that range from conventional numerical error metrics to physically motivated evaluation tasks. With a focus on graph-based MLIP architectures, we then show examples of using this data set to conduct numerical experiments, systematically assessing (i) the effect of hyperparameters on task-level performance and (ii) the fine-tuning behavior of selected pretrained ("foundational") MLIP models. Beyond the Li-P-S solid-state electrolytes, we expect that such benchmarks and accompanying code can be readily adapted to other material systems.
Fragapane et al. (2026) studied this question.