This research paper presents a comparative benchmarking study of modern observability tools including Prometheus, Grafana, Datadog, and OpenTelemetry in high-traffic cloud-native environments. The study evaluates performance, scalability, resource utilization, query latency, metrics ingestion throughput, and distributed tracing overhead under workloads reaching up to 10 million events per second. The research aims to help DevOps and Site Reliability Engineering (SRE) teams select efficient observability solutions for large-scale distributed systems based on operational requirements, scalability needs, and cost-performance trade-offs.
Keshav Prajapati (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: