Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
May 20, 2026Proceedings of the ACM on Management of Data

Efficient LLM Serving for Agentic Workflows: A Data Systems Perspective

View Full Paper
Ask AI
Bookmark
Share

Authors

NWNoppanat WadlomJSJunyi ShenYLYao Lu

Discussion

Loading...

Member takes

Overview

Randomized trial shows up to 1.56× speedup in LLM workflows, suggesting end-to-end optimization is vital for efficiency.

Key Points

  • This research aims to optimize the serving of agentic workflows involving Large Language Models (LLMs) by addressing inefficiencies in current systems.
  • Introduced Helium, a workflow-aware serving framework for LLM invocations.
  • Integrated proactive caching and cache-aware scheduling to enhance prompt reuse.
  • Modelled agentic workloads as query plans to leverage classic query optimization.
  • Achieved up to 1.56× speedup over existing agent serving systems.
  • Demonstrated improved efficiency across various workloads through end-to-end optimization.

Cite This Study

Wadlom et al. (2026) studied this question.

synapsesocial.com/papers/6a0d4f34f03e14405aa9a667https://doi.org/10.1145/3802046
View Full Paper
Ask AI
Bookmark
Share