PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 29, 20250 citationsOpen Access

Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding

View Full Paper
YWYixuan WangSLShuaicheng LiuSJShiyu ji

Key Points

  • Reflective Verification increases the acceptance length of draft tokens, enhancing performance in diverse settings.
  • This method achieves up to a 15% improvement in decoding speed by effectively verifying token semantics.
  • By leveraging LLMs' reflective capacity, verification methods focus on both consistency and correctness.
  • Experiments on multiple benchmarks show significant speedups without compromising model performance.

Abstract

Large language models (LLMs) suffer from high inference latency due to the auto-regressive decoding process. Speculative decoding accelerates inference by generating multiple draft tokens using a lightweight model and verifying them in parallel. However, existing verification methods rely heavily on distributional consistency while overlooking semantic correctness, thereby limiting the potential speedup of speculative decoding. While some methods employ additional models for relaxed verification of draft tokens, they often fail to generalize effectively to more diverse or open-domain settings. In this work, we propose Reflective Verification, a training-free and semantics-aware approach that achieves a better trade-off between correctness and efficiency. Specifically, we leverage the inherent reflective capacity of LLMs to semantically assess the correctness of draft tokens in parallel during verification. Using prompt-based probing, we obtain both the original and reflective distributions of draft tokens in a single forward pass. The fusion of these distributions enables semantic-level verification of draft tokens that incorporates both consistency and correctness. Experiments across multiple domain benchmarks and model scales demonstrate that our method significantly increases the acceptance length of draft tokens without compromising model performance. Furthermore, we find that the proposed Reflective Verification is orthogonal to existing statistical verification methods, and their combination yields additional 515\% improvements in decoding speed.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2025) studied this question.

synapsesocial.com/papers/68da58d8c1728099cfd1118ahttps://doi.org/10.48550/arxiv.2505.18629
Ask AI
Helpful
Bookmark
Share
View Full Paper