conv.

All stories
AIQuiet 47h · day 9

PrismML launches Bonsai 2, a 5.9GB model matching 98% of Qwen3.8 27B's benchmarks

The Caltech-founded startup's ternary-weight compression squeezes a 27B-parameter open model down to smartphone-friendly size with minimal performance loss.

What to know

  • Bonsai 2 27B compresses Alibaba's Qwen3.8 27B from a much larger footprint down to 5.9 GB — a 9x to 10x memory reduction — while retaining about 98% of its benchmark performance.
  • The compression relies on 'ternary' weights (+1, −1, or 0) instead of standard 16-bit values, drastically cutting storage needs.
  • PrismML, founded by Caltech researchers and led by CEO Babak Hassibi, has raised only a $22.25 million seed round but counts Databricks co-founder Ion Stoica as an adviser and is reportedly in talks with Apple.
  • The company's prior model, the original Bonsai released in March, has been downloaded over 11 million times, and PrismML plans to apply the technique to models with several hundred billion parameters within a couple of months.

Babak Hassibi CEO of PrismMLIon Stoica PrismML adviserPrismML AI startup

PrismML launches Bonsai 2, a 5.9GB model matching 98% of Qwen3.8 27B's benchmarks
techcrunch.com

How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts

Peak 5 pieces in two hours at Sep 17, 9 PM; 18 pieces over 9 days (6 articles · 12 posts) Sep 17, 5 PM — 4 pieces · 2 articles · 2 posts — Mastodon 2, Google News 1, Newswires 1Sep 17, 7 PM — quietSep 17, 9 PM — 5 pieces · 2 articles · 3 posts — Newswires 2, Bluesky 1, Mastodon 1, +1 moreSep 17, 11 PM — quietSep 18, 1 AM — quietSep 18, 3 AM — 1 piece · 1 post — Bluesky 1Sep 18, 5 AM — quietSep 18, 7 AM — 2 pieces · 2 posts — Bluesky 1, Reddit 1Sep 18, 9 AM — quietSep 18, 11 AM — 1 piece · 1 article — Google News 1Sep 18, 1 PM — quietSep 18, 3 PM — quietSep 18, 5 PM — quietSep 18, 7 PM — quietSep 18, 9 PM — quietSep 18, 11 PM — 1 piece · 1 post — Mastodon 1Sep 19, 1 AM — quietSep 19, 3 AM — quietSep 19, 5 AM — quietSep 19, 7 AM — quietSep 19, 9 AM — quietSep 19, 11 AM — quietSep 19, 1 PM — quietSep 19, 3 PM — quietSep 19, 5 PM — quietSep 19, 7 PM — quietSep 19, 9 PM — quietSep 19, 11 PM — quietSep 20, 1 AM — quietSep 20, 3 AM — quietSep 20, 5 AM — quietSep 20, 7 AM — quietSep 20, 9 AM — quietSep 20, 11 AM — quietSep 20, 1 PM — quietSep 20, 3 PM — quietSep 20, 5 PM — quietSep 20, 7 PM — 1 piece · 1 post — Bluesky 1Sep 20, 9 PM — quietSep 20, 11 PM — quietSep 21, 1 AM — quietSep 21, 3 AM — quietSep 21, 5 AM — quietSep 21, 7 AM — quietSep 21, 9 AM — quietSep 21, 11 AM — quietSep 21, 1 PM — 1 piece · 1 post — Bluesky 1Sep 21, 3 PM — quietSep 21, 5 PM — quietSep 21, 7 PM — quietSep 21, 9 PM — quietSep 21, 11 PM — quietSep 22, 1 AM — quietSep 22, 3 AM — quietSep 22, 5 AM — quietSep 22, 7 AM — quietSep 22, 9 AM — quietSep 22, 11 AM — quietSep 22, 1 PM — quietSep 22, 3 PM — quietSep 22, 5 PM — quietSep 22, 7 PM — quietSep 22, 9 PM — quietSep 22, 11 PM — quietSep 23, 1 AM — quietSep 23, 3 AM — quietSep 23, 5 AM — quietSep 23, 7 AM — quietSep 23, 9 AM — quietSep 23, 11 AM — quietSep 23, 1 PM — quietSep 23, 3 PM — quietSep 23, 5 PM — quietSep 23, 7 PM — quietSep 23, 9 PM — quietSep 23, 11 PM — quietSep 24, 1 AM — quietSep 24, 3 AM — quietSep 24, 5 AM — quietSep 24, 7 AM — quietSep 24, 9 AM — quietSep 24, 11 AM — quietSep 24, 1 PM — 2 pieces · 1 article · 1 post — Mastodon 1, Newswires 1Sep 24, 3 PM — quietSep 24, 5 PM — quietSep 24, 7 PM — quietSep 24, 9 PM — quietSep 24, 11 PM — quietYesterday, 1 AM — quietYesterday, 3 AM — quietYesterday, 5 AM — quietYesterday, 7 AM — quietYesterday, 9 AM — quietYesterday, 11 AM — quietYesterday, 1 PM — quietYesterday, 3 PM — quietYesterday, 5 PM — quietYesterday, 7 PM — quietYesterday, 9 PM — quietYesterday, 11 PM — quietToday, 1 AM — quietToday, 3 AM — quietToday, 5 AM — quietToday, 7 AM — quietToday, 9 AM — quietToday, 11 AM — quietToday, 1 PM — quiet 1–2
Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24yesterdaynow · 2:50 PM ET
  1. 1

    Early users praise the compressed model as a practical local-AI option

    A Bluesky user highlighted the model's usefulness for recommending capable local AI to people with ordinary laptops, pointing to the 6 GB weights and a browser-native WebGPU demo.

    “it's so nice to have a capable local model i can recommend to people in my life who only have normal laptops, super valuable drop.”
    — @crumb
    1. first by RuntimeWire, 8d ago

      1 more headline
    • it's so nice to have a capable local model i can recommend to people in my life who only have normal laptops, super valuable drop. check out the new prism-ml quantization of Qwen3.8-27b. 6GB weights! — blog: prismml.com/news/bonsai-... webgpu (browser-native!) demo: huggingface.co/spaces/webml... …

      @crumbBluesky8d agoview on Bluesky ↗
    2 more of the top 3 · 5 posts in this stretch
    • everyone is obsessed with massive models but prismml gets it. tiny llms will win because latency is the only feature that actually matters for ai video. the era of waiting for a token stream to finish is dead.

      @xcelestiusX8d agoview on X ↗
    • vick21@mastodon.social

      Wow, this is a true wow! Installing now… https:// techcrunch.com/2026/09/17/pris mml-hopes-its-tiny-llm-could-change-how-we-all-use-ai/

      vick21@mastodon.socialMastodon7d agoview on Mastodon ↗
    all of them →
  2. 2

    Techmeme and outlets pick up the Bonsai 2 release

    Aggregators and social accounts amplified the TechCrunch report, citing the 5.9 GB size and 98.2% benchmark retention figure.

    “The next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I e…”
    — Babak Hassibi, CEO of PrismML · source
  3. background

    PrismML releases Bonsai 2 27B, compressing Qwen3.8 27B to 5.9 GB — The new model uses PrismML's 'ternary' weight compression (+1, −1, or 0 instead of 16-bit values) to shrink Alibaba's Qwen3.8 27B by 9x to 10x, down to 5.9 GB — small enough for a PC or high-end smartphone — while retaining about 98% of Qwen's aggregate benchmark scores.

  4. background

    PrismML releases first Bonsai model at 95% benchmark parity — PrismML's original Bonsai compressed model matched 95% of its source model's aggregate benchmark scores and went on to be downloaded over 11 million times.

Also covered reported alongside — the timeline has no entry for these yet

  1. first by TechCrunch, 8d ago

    1 more headline

and 2 smaller pieces

What people are saying 2 voices from 1 site · best of 5 · verbatim