Two arXiv papers examine efficiency of speculative decoding techniques
Researchers question whether recent shortcuts to speed up large language models actually deliver on their promises.
What to know
- Two new arXiv papers challenge the efficiency of popular speculative decoding methods used to accelerate large language model inference.
- One identifies wasted computation from rejected tokens; another questions whether a hybrid approach actually achieves lossless acceleration under real numerical precision conditions.
- A practical explainer from an inference engineer walks through when and how speculative decoding saves time in production LLM serving.
“By construction, verification computes representations for both accepted and rejected tokens. Yet, conventional drafters retain only the representations of accepted tokens, leaving the substantial verifier computation spent on rejected tokens effectively wasted.”
Koo et al., Researchers · arXiv ↗ · Sep 14
Jahyun Koo, Sunghyeon Woo, Jaeeun Kil, Jeongtae Lee, Sungjae Lee, Kyomin Jung, Minsub Kim ResearchersIlya Koziev, Leonid Sinev, Ivan Oseledets ResearchersDanial Hasan Inference engineer
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
-
2
Hasan publishes inference engineering explainer on speculative decoding
Engineer Danial Hasan publishes an article explaining speculative decoding as part of an inference engineering series, detailing how the technique speeds LLM inference and when it saves time, with practical examples using open-source vLLM.
“Speculative decoding can reduce the time an AI assistant takes to finish an answer. A cheap proposer drafts several tokens, and the model serving the request checks them together.”
— Danial Hasan -
article 3 of my inference engineering series - diving into 1 of the 4 major techniques, speculative decoding! this technique speeds up LLM inference by using a smaller, faster model to draft tokens that the larger model verifies in parallel.
-
-
1
Koziev et al. question Orthrus's lossless decoding claim
A second arXiv paper independently reproduces the Orthrus hybrid architecture and examines whether its central claim—that an intra-model consensus mechanism enables lossless speculative decoding—holds under different numerical precision conditions.
-
first by arXiv cs.AI, 12d ago
-
-
background
Koo et al. propose carryover drafting to recycle rejected tokens — Researchers identify that conventional speculative decoding discards representations of rejected tokens despite the verifier computing them, proposing a method called "carryover drafting" to retain and reuse this wasted computation.