PrismML launches Bonsai 2, a 5.9GB model matching 98% of Qwen3.8 27B's benchmarks
The Caltech-founded startup's ternary-weight compression squeezes a 27B-parameter open model down to smartphone-friendly size with minimal performance loss.
What to know
- Bonsai 2 27B compresses Alibaba's Qwen3.8 27B from a much larger footprint down to 5.9 GB — a 9x to 10x memory reduction — while retaining about 98% of its benchmark performance.
- The compression relies on 'ternary' weights (+1, −1, or 0) instead of standard 16-bit values, drastically cutting storage needs.
- PrismML, founded by Caltech researchers and led by CEO Babak Hassibi, has raised only a $22.25 million seed round but counts Databricks co-founder Ion Stoica as an adviser and is reportedly in talks with Apple.
- The company's prior model, the original Bonsai released in March, has been downloaded over 11 million times, and PrismML plans to apply the technique to models with several hundred billion parameters within a couple of months.
Babak Hassibi CEO of PrismMLIon Stoica PrismML adviserPrismML AI startup
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
-
1
Early users praise the compressed model as a practical local-AI option
A Bluesky user highlighted the model's usefulness for recommending capable local AI to people with ordinary laptops, pointing to the 6 GB weights and a browser-native WebGPU demo.
“it's so nice to have a capable local model i can recommend to people in my life who only have normal laptops, super valuable drop.”
— @crumb -
first by RuntimeWire, 8d ago
-
it's so nice to have a capable local model i can recommend to people in my life who only have normal laptops, super valuable drop. check out the new prism-ml quantization of Qwen3.8-27b. 6GB weights! — blog: prismml.com/news/bonsai-... webgpu (browser-native!) demo: huggingface.co/spaces/webml... …
2 more of the top 3 · 5 posts in this stretch
-
everyone is obsessed with massive models but prismml gets it. tiny llms will win because latency is the only feature that actually matters for ai video. the era of waiting for a token stream to finish is dead.
-
V
Wow, this is a true wow! Installing now… https:// techcrunch.com/2026/09/17/pris mml-hopes-its-tiny-llm-could-change-how-we-all-use-ai/
-
-
2
Techmeme and outlets pick up the Bonsai 2 release
Aggregators and social accounts amplified the TechCrunch report, citing the 5.9 GB size and 98.2% benchmark retention figure.
“The next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I e…”
— Babak Hassibi, CEO of PrismML · source -
background
PrismML releases Bonsai 2 27B, compressing Qwen3.8 27B to 5.9 GB — The new model uses PrismML's 'ternary' weight compression (+1, −1, or 0 instead of 16-bit values) to shrink Alibaba's Qwen3.8 27B by 9x to 10x, down to 5.9 GB — small enough for a PC or high-end smartphone — while retaining about 98% of Qwen's aggregate benchmark scores.
-
background
PrismML releases first Bonsai model at 95% benchmark parity — PrismML's original Bonsai compressed model matched 95% of its source model's aggregate benchmark scores and went on to be downloaded over 11 million times.
Also covered reported alongside — the timeline has no entry for these yet
-
first by TechCrunch, 8d ago
1 more headline
- PrismML hopes its tiny LLM could change how we all use AI TechCrunch · 8d ago
and 2 smaller pieces
What people are saying 2 voices from 1 site · best of 5 · verbatim
- Sep 20
-
K
PrismML espera que su pequeño LLM cambie la forma en que todos usamos la IA —
- Sep 18
-
O
If AI lab PrismML isn't on your radar yet, it should be. #Techcrunch #breakingnews #news #othur