ByteShape adds Prism-ML Ternary Bonsai 2 to benchmarks
3 Sep 18 · 8d ago · 2 comments · 1 source · development 3 of 3
ByteShape updated their benchmark post to include Prism-ML's Ternary Bonsai 2 models, which are the fastest points on every GPU tested at 1.77 and 2.14 BPW, though they score 91.4% and 91.7% of BF16 and require custom llama.cpp builds.
“Also model with their draft answered incorrectly. With MTP it answered correctly.”
npodbielski, Hacker News commenter, AMD 7900XTX user · hn ↗ByteShape Team Quantization developerAlibaba Model creator
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 2 voices · verbatim
-
Well I tested it on 7900XTX with the same prompts and their draft model gave me about 30t/s. Their own snippet of code with regular MTP model gave me 60t/s.Also model with their draft answered incorrectly. With MTP it answered correctly.Question was: "Does MikroTik CRS312-4C+8XG-RM have combo ports?". The answer is Yes.
-
I’m on a strix halo @ GPU-5 with MTP and I get 600 prefill and 30 TG which pushes it into a very usable range. The odd thing is that Dflash2 is really slow for me, like sub 10 TG.
All 3 developments of ByteShape releases full ShapeLearn Qwen 3.8 27B… →
Hacker NewsMastodonNewswires