Erasure-coded storage systems based on Ceph have become a mainstay within UK Grid sites as a means of providing bulk data storage whilst maintaining a good balance between data safety and space efficiency. These storage systems are complex and self-correcting, but despite access to a myriad of metrics, the inner workings of the storage tend to be opaque to the storage admin. One of the common problems seen within Ceph based systems is slow ops—instances of operations that take longer than expected, that are also often blocking in nature, impacting the overall performance and reliability of the system. Identifying the causes of slow ops can help to prevent or reduce the impact of future occurrences, leading to an increase in performance and reliability. We detail the efforts of the Lancaster Grid Site to understand the causes of and mitigate against these slow ops and other performance bottlenecks within our storage system. We endeavour to bring together a holistic monitoring model, utilising Ceph metrics, detailed XRootD monitoring streams, and client-side logging, in order to understand how data-management events impact the health of the storage.
Doidge et al. (2025) studied this question.