This paper proposes CAIR (Carbon-Aware Inference Router), a real-time routing framework for large language models. Requests are routed between model tiers based on task complexity and live grid carbon intensity, targeting measurable emissions reduction without accuracy or latency loss. Preliminary analysis on a 1M prompt/day system suggests ~62% reduction in inference carbon. Framework repository: https://github.com/pretzelslab/sa1-carbon-inference-router
Preethi Raghuveeran (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: