conv.

All stories
AIQuiet 12d · day 13

Two arXiv papers examine efficiency of speculative decoding techniques

Researchers question whether recent shortcuts to speed up large language models actually deliver on their promises.

What to know

  • Two new arXiv papers challenge the efficiency of popular speculative decoding methods used to accelerate large language model inference.
  • One identifies wasted computation from rejected tokens; another questions whether a hybrid approach actually achieves lossless acceleration under real numerical precision conditions.
  • A practical explainer from an inference engineer walks through when and how speculative decoding saves time in production LLM serving.

“By construction, verification computes representations for both accepted and rejected tokens. Yet, conventional drafters retain only the representations of accepted tokens, leaving the substantial verifier computation spent on rejected tokens effectively wasted.”

Koo et al., Researchers · arXiv ↗ · Sep 14

Jahyun Koo, Sunghyeon Woo, Jaeeun Kil, Jeongtae Lee, Sungjae Lee, Kyomin Jung, Minsub Kim ResearchersIlya Koziev, Leonid Sinev, Ivan Oseledets ResearchersDanial Hasan Inference engineer

Two arXiv papers examine efficiency of speculative decoding techniques
x.com

How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts

Peak 2 pieces in 3h at Sep 14, 11 PM; 3 pieces over 13 days (2 articles · 1 post) Sep 14, 11 PM — 2 pieces · 2 articles — Newswires 2Sep 15, 2 AM — quietSep 15, 5 AM — quietSep 15, 8 AM — quietSep 15, 11 AM — quietSep 15, 2 PM — 1 piece · 1 post — X 1Sep 15, 5 PM — quietSep 15, 8 PM — quietSep 15, 11 PM — quietSep 16, 2 AM — quietSep 16, 5 AM — quietSep 16, 8 AM — quietSep 16, 11 AM — quietSep 16, 2 PM — quietSep 16, 5 PM — quietSep 16, 8 PM — quietSep 16, 11 PM — quietSep 17, 2 AM — quietSep 17, 5 AM — quietSep 17, 8 AM — quietSep 17, 11 AM — quietSep 17, 2 PM — quietSep 17, 5 PM — quietSep 17, 8 PM — quietSep 17, 11 PM — quietSep 18, 2 AM — quietSep 18, 5 AM — quietSep 18, 8 AM — quietSep 18, 11 AM — quietSep 18, 2 PM — quietSep 18, 5 PM — quietSep 18, 8 PM — quietSep 18, 11 PM — quietSep 19, 2 AM — quietSep 19, 5 AM — quietSep 19, 8 AM — quietSep 19, 11 AM — quietSep 19, 2 PM — quietSep 19, 5 PM — quietSep 19, 8 PM — quietSep 19, 11 PM — quietSep 20, 2 AM — quietSep 20, 5 AM — quietSep 20, 8 AM — quietSep 20, 11 AM — quietSep 20, 2 PM — quietSep 20, 5 PM — quietSep 20, 8 PM — quietSep 20, 11 PM — quietSep 21, 2 AM — quietSep 21, 5 AM — quietSep 21, 8 AM — quietSep 21, 11 AM — quietSep 21, 2 PM — quietSep 21, 5 PM — quietSep 21, 8 PM — quietSep 21, 11 PM — quietSep 22, 2 AM — quietSep 22, 5 AM — quietSep 22, 8 AM — quietSep 22, 11 AM — quietSep 22, 2 PM — quietSep 22, 5 PM — quietSep 22, 8 PM — quietSep 22, 11 PM — quietSep 23, 2 AM — quietSep 23, 5 AM — quietSep 23, 8 AM — quietSep 23, 11 AM — quietSep 23, 2 PM — quietSep 23, 5 PM — quietSep 23, 8 PM — quietSep 23, 11 PM — quietSep 24, 2 AM — quietSep 24, 5 AM — quietSep 24, 8 AM — quietSep 24, 11 AM — quietSep 24, 2 PM — quietSep 24, 5 PM — quietSep 24, 8 PM — quietSep 24, 11 PM — quietSep 25, 2 AM — quietSep 25, 5 AM — quietSep 25, 8 AM — quietSep 25, 11 AM — quietSep 25, 2 PM — quietSep 25, 5 PM — quietSep 25, 8 PM — quietSep 25, 11 PM — quietYesterday, 2 AM — quietYesterday, 5 AM — quietYesterday, 8 AM — quietYesterday, 11 AM — quietYesterday, 2 PM — quietYesterday, 5 PM — quietYesterday, 8 PM — quietYesterday, 11 PM — quietToday, 2 AM — quietToday, 5 AM — quietToday, 8 AM — quietToday, 11 AM — quietToday, 2 PM — quietToday, 5 PM — quietToday, 8 PM — quiet 12
Sep 15Sep 16Sep 17Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25yesterdaynow · 11:43 PM ET
  1. 2

    Hasan publishes inference engineering explainer on speculative decoding

    Engineer Danial Hasan publishes an article explaining speculative decoding as part of an inference engineering series, detailing how the technique speeds LLM inference and when it saves time, with practical examples using open-source vLLM.

    “Speculative decoding can reduce the time an AI assistant takes to finish an answer. A cheap proposer drafts several tokens, and the model serving the request checks them together.”
    — Danial Hasan
    • article 3 of my inference engineering series - diving into 1 of the 4 major techniques, speculative decoding! this technique speeds up LLM inference by using a smaller, faster model to draft tokens that the larger model verifies in parallel.

      @danialhasanX12d ago43▲view on X ↗
  2. 1

    Koziev et al. question Orthrus's lossless decoding claim

    A second arXiv paper independently reproduces the Orthrus hybrid architecture and examines whether its central claim—that an intra-model consensus mechanism enables lossless speculative decoding—holds under different numerical precision conditions.

    1. first by arXiv cs.AI, 12d ago

  3. background

    Koo et al. propose carryover drafting to recycle rejected tokens — Researchers identify that conventional speculative decoding discards representations of rejected tokens despite the verifier computing them, proposing a method called "carryover drafting" to retain and reuse this wasted computation.