conv.

All stories
AIQuiet 8d · day 9

ByteShape releases full ShapeLearn Qwen 3.8 27B quantizations

ByteShape published optimized quantizations of Qwen 3.8 27B, achieving 99.63% accuracy of the original model with gains in throughput across multiple GPUs.

What to know

  • ByteShape's ShapeLearn full quantizations of Qwen 3.8 27B achieve 99.63% accuracy of the original model with GPU-5 reaching 93.7 tokens/second throughput.
  • Performance gains vary significantly by hardware: Nvidia GPUs show strong results, while AMD users report mixed outcomes with some finding Unsloth models competitive or faster.
  • Memory bandwidth emerges as the primary performance bottleneck, explaining why higher-bandwidth cards (4090, 5090) see larger benefits from quantization optimizations than lower-bandwidth AMD cards.
  • The draft model (DFlash2) shows accuracy concerns in user testing despite speed gains, raising tradeoffs between throughput and correctness.

The dispute Whether ShapeLearn quantizations deliver meaningful real-world performance gains: Nvidia GPU users and those focused on accurate inference see clear wins, while AMD users and those testing draft modes report they match or underperform existing alternatives. · positions read across 8 posts and comments

many voices

Benchmarks are valuable but show limited real-world gains on AMD/lower-bandwidth hardware.

some voices

Draft models trade accuracy for speed, making faster modes unsuitable for reliable inference.

  • “Also model with their draft answered incorrectly. With MTP it answered correctly.”

    npodbielski · Hacker News ↗
some voices

The benchmarking methodology and GPU comparisons are exceptionally useful for optimization choices.

  • “Absolute treasure of a website with these graphs, thank you for sharing this. Huge help for me to find a faster model (smaller quantization) for my VRAM.”

    Schlagbohrer · Hacker News ↗

ByteShape Team Quantization developerAlibaba Model creator

ByteShape releases full ShapeLearn Qwen 3.8 27B quantizations
byteshape.com

How it unfolded 3 developments, newest first · click a bar or a number to jump articlespostscomments

Peak 3 pieces in two hours at Sep 17, 11 PM; 11 pieces over 9 days (1 article · 2 posts · 8 comments) Sep 17, 9 PM — 2 pieces · 1 article · 1 post — Hacker News 1, Newswires 1Sep 17, 11 PM — 3 pieces · 1 post · 2 comments — Hacker News 2, Mastodon 1Sep 18, 1 AM — quietSep 18, 3 AM — 2 pieces · 2 comments — Hacker News 2Sep 18, 5 AM — 2 pieces · 2 comments — Hacker News 2Sep 18, 7 AM — 2 pieces · 2 comments — Hacker News 2Sep 18, 9 AM — quietSep 18, 11 AM — quietSep 18, 1 PM — quietSep 18, 3 PM — quietSep 18, 5 PM — quietSep 18, 7 PM — quietSep 18, 9 PM — quietSep 18, 11 PM — quietSep 19, 1 AM — quietSep 19, 3 AM — quietSep 19, 5 AM — quietSep 19, 7 AM — quietSep 19, 9 AM — quietSep 19, 11 AM — quietSep 19, 1 PM — quietSep 19, 3 PM — quietSep 19, 5 PM — quietSep 19, 7 PM — quietSep 19, 9 PM — quietSep 19, 11 PM — quietSep 20, 1 AM — quietSep 20, 3 AM — quietSep 20, 5 AM — quietSep 20, 7 AM — quietSep 20, 9 AM — quietSep 20, 11 AM — quietSep 20, 1 PM — quietSep 20, 3 PM — quietSep 20, 5 PM — quietSep 20, 7 PM — quietSep 20, 9 PM — quietSep 20, 11 PM — quietSep 21, 1 AM — quietSep 21, 3 AM — quietSep 21, 5 AM — quietSep 21, 7 AM — quietSep 21, 9 AM — quietSep 21, 11 AM — quietSep 21, 1 PM — quietSep 21, 3 PM — quietSep 21, 5 PM — quietSep 21, 7 PM — quietSep 21, 9 PM — quietSep 21, 11 PM — quietSep 22, 1 AM — quietSep 22, 3 AM — quietSep 22, 5 AM — quietSep 22, 7 AM — quietSep 22, 9 AM — quietSep 22, 11 AM — quietSep 22, 1 PM — quietSep 22, 3 PM — quietSep 22, 5 PM — quietSep 22, 7 PM — quietSep 22, 9 PM — quietSep 22, 11 PM — quietSep 23, 1 AM — quietSep 23, 3 AM — quietSep 23, 5 AM — quietSep 23, 7 AM — quietSep 23, 9 AM — quietSep 23, 11 AM — quietSep 23, 1 PM — quietSep 23, 3 PM — quietSep 23, 5 PM — quietSep 23, 7 PM — quietSep 23, 9 PM — quietSep 23, 11 PM — quietSep 24, 1 AM — quietSep 24, 3 AM — quietSep 24, 5 AM — quietSep 24, 7 AM — quietSep 24, 9 AM — quietSep 24, 11 AM — quietSep 24, 1 PM — quietSep 24, 3 PM — quietSep 24, 5 PM — quietSep 24, 7 PM — quietSep 24, 9 PM — quietSep 24, 11 PM — quietYesterday, 1 AM — quietYesterday, 3 AM — quietYesterday, 5 AM — quietYesterday, 7 AM — quietYesterday, 9 AM — quietYesterday, 11 AM — quietYesterday, 1 PM — quietYesterday, 3 PM — quietYesterday, 5 PM — quietYesterday, 7 PM — quietYesterday, 9 PM — quietYesterday, 11 PM — quietToday, 1 AM — quietToday, 3 AM — quietToday, 5 AM — quietToday, 7 AM — quietToday, 9 AM — quietToday, 11 AM — quietToday, 1 PM — quiet 1–3
Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24yesterdaynow · 2:50 PM ET
  1. 3

    ByteShape adds Prism-ML Ternary Bonsai 2 to benchmarks

    ByteShape updated their benchmark post to include Prism-ML's Ternary Bonsai 2 models, which are the fastest points on every GPU tested at 1.77 and 2.14 BPW, though they score 91.4% and 91.7% of BF16 and require custom llama.cpp builds.

    “Also model with their draft answered incorrectly. With MTP it answered correctly.”
    — npodbielski, Hacker News commenter, AMD 7900XTX user · source
    • Well I tested it on 7900XTX with the same prompts and their draft model gave me about 30t/s. Their own snippet of code with regular MTP model gave me 60t/s.Also model with their draft answered incorrectly. With MTP it answered correctly.Question was: "Does MikroTik CRS312-4C+8XG-RM have combo ports?". The answer is Yes.

      npodbielskiHacker News8d agoview on Hacker News ↗
    1 more of the top 2 · 2 posts in this stretch
    • I’m on a strix halo @ GPU-5 with MTP and I get 600 prefill and 30 TG which pushes it into a very usable range. The odd thing is that Dflash2 is really slow for me, like sub 10 TG.

      syntaxingHacker News8d agoview on Hacker News ↗
    all of them →
  2. 2

    Users confirm memory bandwidth as performance bottleneck

    AMD GPU users attribute performance constraints to memory bandwidth limitations, noting the 7900 XTX has similar bandwidth to the 3090 and that Nvidia 4090/5090 cards with higher bandwidth would likely see larger gains from ByteShape's optimizations.

    “I assume the limit for me is memory bandwith, as the 7900 XTX has the same bandwith as the 3090 from what I can gather and I already reached ~60 t/s with Unsloth.”
    — Systemerror7A69
    • AMD 7900 XTX with Vulkan here as well, wasn't faster on my test either. Might be much different on Nvidia though.I assume the limit for me is memory bandwith, as the 7900 XTX has the same bandwith as the 3090 from what I can gather and I already reached ~60 t/s with Unsloth. Those would fit with the numbers Byteshape has for their cards.4090 and…

      Systemerror7A69Hacker News8d agoview on Hacker News ↗
    2 more of the top 3 · 4 posts in this stretch
    • Absolute treasure of a website with these graphs, thank you for sharing this. Huge help for me to find a faster model (smaller quantization) for my VRAM.

      SchlagbohrerHacker News8d agoview on Hacker News ↗
    • I feel weird that I like your typos, because clearly AI did not write your post. Typos have become downright charming and nostalgic for me.

      SchlagbohrerHacker News8d agoview on Hacker News ↗
    all of them →
  3. 1

    Users report mixed performance results on AMD hardware

    Comments from Hacker News users reveal performance varies significantly across hardware. Some AMD GC Vulkan users report ShapeLearn is not faster than Unsloth models, while an AMD 7900 XTX user found draft models gave ~30 tokens/second versus 60 tokens/second with MTP, with accuracy concerns on the draft model.

    “It's not faster than the unsloth model.”
    — _ache_
    • From my own test. It's not faster than the unsloth model.Disclarer: I'm unsing Vulkan on an AMD GC.

      _ache_Hacker News8d agoview on Hacker News ↗
    2 posts in this stretch →
  4. background

    ByteShape publishes full ShapeLearn benchmarks — ByteShape published benchmarks comparing full ShapeLearn models with ShapeLearn-Lite and competing quantizations. All five ShapeLearn models remain on the measured frontier, with GPU-5 achieving the highest aggregate score among plotted quants and reaching 93.7 tokens/second at 99.63% of BF16 baseline.

  5. background

    ByteShape publishes ShapeLearn-Lite quantizations — Four days after Qwen 3.8 27B's release, ByteShape published their first set of GGUFs called ShapeLearn-Lite, produced with smaller optimization budgets and fewer checks than their full release would require.

  6. background

    Alibaba releases Qwen 3.8 27B model — Qwen 3.8 27B was released on August 14, 2026.

What people are saying 1 voices from 1 site · best of 8 · verbatim

Still unanswered
  • Why does DFlash2 draft mode generate incorrect answers compared to the full MTP model?
  • How much does memory bandwidth difference between GPUs account for the performance variance users are seeing?