Authors
Loading...
This framework shows that uniform KV cache quantization is suboptimal in reasoning models, implying better performance with tailored approaches.
Raviteja Nekkalapu (2026) studied this question.