conv.

All stories
AIQuiet 7d · day 9

PrismML releases Ternary Bonsai 2 27B, 9x compressed model retaining 98% performance

New quantized model runs 27B-class AI on consumer hardware with near-lossless capability preservation.

What to know

  • PrismML's Ternary Bonsai 2 27B achieves 9x model compression while retaining 98.2% of benchmark performance—particularly important for coding, tool use, and multimodal tasks where small errors compound.
  • The 5.9GB model runs on consumer hardware (RTX 4090, Mac M5 Max) and in browsers via WebGPU, enabling local AI deployment without cloud dependency.
  • Community testing reveals genuine capability gains over the previous Bonsai generation but also confirms severe degradation on longer tasks—marketing claims of 'near-lossless' performance face scrutiny.
  • Unclear technical advantage: commenters question how this ternary approach meaningfully differs from existing Q2-level quantization of the same Qwen base model.

The dispute Whether 'near-lossless' performance claims hold up: some users report failure modes on longer sequences, conflicting with the headline 98.2% aggregate benchmark retention figure. · positions read across 42 posts and comments

many voices

This is a genuine advance for local inference, especially for coding work on consumer hardware.

  • “Running at about 7-8 tok/s (~60 tok/s prefill) on a Mac Mini M2 16GB. So far feels smarter than Bonsai 1 27B, it's slightly larger than the Q1_0 quant. Super exciting stuff :)”

    jedbrooke · Hacker News ↗
many voices

The model degrades badly on longer tasks and the marketing 'near-lossless' claim is misleading.

  • “Use it for any longer task and they fall apart spectacularly and in interesting ways.”

    Aurornis · Hacker News ↗
some voices

The technical novelty is unclear; comparisons to standard quantization methods (Q2) and the special sauce are missing from documentation.

  • “I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?”

    adrian17 · Hacker News ↗

PrismML Model developer and publisherAurornis Community tester

PrismML releases Ternary Bonsai 2 27B, 9x compressed model retaining 98% performance
prismml.com

How it unfolded 4 developments, newest first · click a bar or a number to jump articlespostscomments

Peak 10 pieces in two hours at Sep 17, 4 PM; 50 pieces over 9 days (5 articles · 3 posts · 42 comments) Sep 17, 4 PM — 10 pieces · 3 articles · 1 post · 6 comments — Hacker News 7, Newswires 3Sep 17, 6 PM — 4 pieces · 4 comments — Hacker News 4Sep 17, 8 PM — 5 pieces · 1 post · 4 comments — Hacker News 4, Mastodon 1Sep 17, 10 PM — 5 pieces · 2 articles · 3 comments — Hacker News 3, Newswires 2Sep 18, 12 AM — quietSep 18, 2 AM — 4 pieces · 1 post · 3 comments — Hacker News 2, Lobsters 2Sep 18, 4 AM — 3 pieces · 3 comments — Hacker News 3Sep 18, 6 AM — 2 pieces · 2 comments — Lobsters 1, Hacker News 1Sep 18, 8 AM — 3 pieces · 3 comments — Hacker News 3Sep 18, 10 AM — 1 piece · 1 comment — Lobsters 1Sep 18, 12 PM — 1 piece · 1 comment — Lobsters 1Sep 18, 2 PM — 2 pieces · 2 comments — Hacker News 1, Lobsters 1Sep 18, 4 PM — 1 piece · 1 comment — Lobsters 1Sep 18, 6 PM — quietSep 18, 8 PM — 1 piece · 1 comment — Lobsters 1Sep 18, 10 PM — 2 pieces · 2 comments — Lobsters 2Sep 19, 12 AM — 1 piece · 1 comment — Lobsters 1Sep 19, 2 AM — 3 pieces · 3 comments — Lobsters 2, Hacker News 1Sep 19, 4 AM — quietSep 19, 6 AM — quietSep 19, 8 AM — quietSep 19, 10 AM — quietSep 19, 12 PM — 1 piece · 1 comment — Lobsters 1Sep 19, 2 PM — quietSep 19, 4 PM — 1 piece · 1 comment — Hacker News 1Sep 19, 6 PM — quietSep 19, 8 PM — quietSep 19, 10 PM — quietSep 20, 12 AM — quietSep 20, 2 AM — quietSep 20, 4 AM — quietSep 20, 6 AM — quietSep 20, 8 AM — quietSep 20, 10 AM — quietSep 20, 12 PM — quietSep 20, 2 PM — quietSep 20, 4 PM — quietSep 20, 6 PM — quietSep 20, 8 PM — quietSep 20, 10 PM — quietSep 21, 12 AM — quietSep 21, 2 AM — quietSep 21, 4 AM — quietSep 21, 6 AM — quietSep 21, 8 AM — quietSep 21, 10 AM — quietSep 21, 12 PM — quietSep 21, 2 PM — quietSep 21, 4 PM — quietSep 21, 6 PM — quietSep 21, 8 PM — quietSep 21, 10 PM — quietSep 22, 12 AM — quietSep 22, 2 AM — quietSep 22, 4 AM — quietSep 22, 6 AM — quietSep 22, 8 AM — quietSep 22, 10 AM — quietSep 22, 12 PM — quietSep 22, 2 PM — quietSep 22, 4 PM — quietSep 22, 6 PM — quietSep 22, 8 PM — quietSep 22, 10 PM — quietSep 23, 12 AM — quietSep 23, 2 AM — quietSep 23, 4 AM — quietSep 23, 6 AM — quietSep 23, 8 AM — quietSep 23, 10 AM — quietSep 23, 12 PM — quietSep 23, 2 PM — quietSep 23, 4 PM — quietSep 23, 6 PM — quietSep 23, 8 PM — quietSep 23, 10 PM — quietSep 24, 12 AM — quietSep 24, 2 AM — quietSep 24, 4 AM — quietSep 24, 6 AM — quietSep 24, 8 AM — quietSep 24, 10 AM — quietSep 24, 12 PM — quietSep 24, 2 PM — quietSep 24, 4 PM — quietSep 24, 6 PM — quietSep 24, 8 PM — quietSep 24, 10 PM — quietYesterday, 12 AM — quietYesterday, 2 AM — quietYesterday, 4 AM — quietYesterday, 6 AM — quietYesterday, 8 AM — quietYesterday, 10 AM — quietYesterday, 12 PM — quietYesterday, 2 PM — quietYesterday, 4 PM — quietYesterday, 6 PM — quietYesterday, 8 PM — quietYesterday, 10 PM — quietToday, 12 AM — quietToday, 2 AM — quietToday, 4 AM — quietToday, 6 AM — quietToday, 8 AM — quietToday, 10 AM — quietToday, 12 PM — quiet 1–4
Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24yesterdaynow · 2:50 PM ET
  1. 4

    Discussion centers on benchmark claims and practical use cases

    Commenters questioned the novelty of the compression technique relative to existing quantization methods, noted unclear comparisons to standard quants, and discussed whether the model's real-world value lies primarily in coding tasks versus general-purpose use.

    “Models that can run with good speed on affordable consumer hardware for coding only is the dream.”
    — blactuary
    • Is it just me or are local models improving much more quickly than the frontier models? If so, that would be a very welcome development, with local/private inference becoming available to everyone.

      mpweihervibecoding8d ago6▲view on Lobsters ↗
    2 more of the top 3 · 38 posts in this stretch
    • I managed to run it with RTX 3070 (8GB VRAM) following the "Quickstart" on HF model card with minor modifications (modify the architecture 86 for your own hardware): git clone https://github.com/PrismML-Eng/llama.cpp && cd llama.cpp cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86 && cmake --build build -j Then downloaded & verified…

      wombat23Hacker News8d agoview on Hacker News ↗
    • The fact that you get 98.2% of the benchmark performance after going from 16 bits per weight to 1.76 bits per weight really shows you how much space is wasted in these models. Wow.

      DustyFuzzyvibecoding8d ago4▲view on Lobsters ↗
    all of them →
  2. 3

    Users identify technical limitations and quantization comparisons

    Community members raised questions about model degradation on longer tasks, requested comparisons with standard Q2 quantizations of the same base model, and noted that existing llama.cpp implementations require PrismML's custom fork to support the ternary weights.

    “Use it for any longer task and they fall apart spectacularly and in interesting ways.”
    — Aurornis
    • If you want to try out out the GGUFs from https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#th... be aware that you need Prism's llama.cpp fork to get them to work, from https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-...This should work: cd /tmp # Get the Prism macOS runtime curl -fL…

      simonwHacker News8d agoview on Hacker News ↗
    1 more of the top 2 · 2 posts in this stretch
    • I really wish people would stop saying N times smaller than something when making a comparison; that makes no sense - it's 1/9th (11.11%) the size. You don't get a smaller quantity by multiplying by a number greater than 1.0. You could instead reverse the subjects being compared - "the original model is 9x bigger than this new smaller, efficient…

      miffy900Hacker News8d agoview on Hacker News ↗
    all of them →
  3. 2

    Community explores model deployment in browsers and on local hardware

    Developers report successfully running the model in browsers via WebGPU kernels and on consumer hardware including Mac Mini M2 and NVIDIA GPUs, with observed throughputs of 7-8 tokens/second on M2 and 143 tokens/second on RTX 5090.

    “These are small enough that you can run them entirely in the browser…”
    — Aurornis
    • These are small enough that you can run them entirely in the browser https://huggingface.co/spaces/webml-community/ternary-bonsai...Remember to clear the downloaded weights afterward.Like the last model, it's amazing they work as well as they do. Use it for any longer task and they fall apart spectacularly and in interesting ways.

      AurornisHacker News8d agoview on Hacker News ↗
  4. 1

    PrismML releases Ternary Bonsai 2 27B model

    PrismML announced Ternary Bonsai 2 27B, a quantized 27B-parameter model using ternary {−1, 0, +1} weights with FP16 group-wise scaling. The model achieves a 5.9GB footprint (1.76 effective bits per weight), more than 9x smaller than full-precision Qwen3.8 27B, while retaining 98.2% of aggregate benchmark performance across reasoning, math, coding, vision, and agentic tool use.

    “Against its full-precision counterpart, Ternary Bonsai 2 27B is more than 9x smaller while retaining 98.2% of aggregate benchmark performance.”
    — PrismML
    1. first by Lobsters, 8d ago · also PrismML News, Simon Willison, HN Best, HN Frontpage

      1 more headline
    • Love this for the folks with 16gb graphics cards - 3.8 27b has been incredible but not quite runnable on anything less than 32gb - will try loading this up on my 16gb intel b50 and see how it goes - not sure these quants can be accelerated by the XPU cores yet but maybe in time!

      kamranjonHacker News8d agoview on Hacker News ↗

What people are saying 18 voices from 2 sites · best of 42 · verbatim

Still unanswered
  • How does Ternary Bonsai 2 actually compare to standard Q2 quantizations of the same Qwen model in real-world coding tasks?
  • Why does the model fail catastrophically on longer tasks if it truly retains 98.2% of performance?
  • What is the 'special sauce' of ternary quantization that distinguishes it from existing low-bit quantization approaches?