Unsloth releases Dynamic 3.0 quantization with >10% accuracy gain
New GGUF quantization method delivers better model quality at same file size for Qwen 3.8-27B.
What to know
- Unsloth's Dynamic 3.0 quantization delivers >10% better top-1% accuracy than competitors at identical file sizes using improved calibration data and post-training techniques.
- The smallest variant (UD-IQ1_S) achieves 72% accuracy at 6.2GB—89% smaller than full precision—but real-world coding performance remains unvalidated by users.
- Community feedback highlights practical concerns: file versioning confusion, MTP removal trade-offs, and demand for code-generation benchmarks beyond KL divergence metrics.
The dispute Whether KL Divergence and top-1% accuracy metrics reliably predict practical code generation performance in production settings. · positions read across 16 posts and comments
Published metrics are insufficient without real-world code generation validation.
-
“Low KL divergence does not mean much when the model gets stuck in doom loops all the time.”
johndough · Hacker News
Unsloth quants are high-quality and worth using with proper tooling.
-
“Your gguf are the first ones I look for when I want to download a gguf model.”
xlayn · Hacker News
File management and metadata practices need improvement for practical deployment.
-
“It would be nice if unsloth published GGUFs would use a version number or something, because now I have multiple different files on local storage that otherwise have exactly the same name.”
walrus01 · Hacker News
Unsloth Quantization library developer
The record 2 articles and posts · last 30 days
- summary covers to here · Aug 19, 10:48 PM · 1 piece above arrived after
What people are saying 16 voices from 2 sites · verbatim
- How do the smaller quants (UD-IQ1_S, UD-IQ2_XXS) perform on real code generation tasks with multiple reasoning steps?
- What is the performance impact of removing MTP from smaller quantizations on inference speed for memory-constrained devices?
- Aug 20
-
Will this also be applied to older models like Qwen3.6? 35B-A3B still has its uses with its higher speed than the dense 3.8 27B.
-
These are very good!I'm hoping for speed improvements because the only problem running the 27B model on my Macbook pro (M4 Max) is the speed: 20 tokens per second. I benchmarked and MTP actually makes things slower, so I disabled MTP altogether. I'm hoping there will be some breakthroughs or optimizations that will allow me to run this at 30-50…
-
N
Unsloth Dynamic 3.0 GGUFs: https:// unsloth.ai/docs/basics/dynamic -3.0-ggufs Discussion: http:// news.ycombinator.com/item?id=4 9365443
-
H
Unsloth Dynamic 3.0 GGUFs Link: https:// unsloth.ai/docs/basics/dynamic -3.0-ggufs Discussion: https:// news.ycombinator.com/item?id=4 9365443
- Aug 19
-
huh, sounds like they’re talking about over fitting and datasets etc, it seems like this is almost more like a fine tune/distill than just a pure quantization
-
You can manage these easily with the huggingface python app.hf download hf://unsloth/Qwen3.8-27B-GGUF \ Qwen3.8-27B-UD-Q4_K_XL.ggufand then see them with `hf cache ls`.Prune old versions with `hf cache prune`.
-
I wonder if there should be some kind of header sort of like a README inside mixed with something like EXIF information that camera images embed.
-
No Dynamic 3.0 NVFP4 quants just yet from the look of it, as a heads up. Would love to see how those perform relative to others on the curve.
-
Since it seems like this not only improved sizes but also performance I can't wait for some benchmarks and comparisons. If you don't have a separate GPU for inference, every single GB matters so a comparison between specific Q4 Quants is really interesting to me.Currently I very much can't decide between going for a bit of a lower Q4 Quant to…
-
It would be nice if unsloth published GGUFs would use a version number or something, because now I have multiple different files on local storage that otherwise have exactly the same name."Qwen3.8-27B-UD-Q8_K_XL.gguf" for instance.The one downloaded at least 4 days ago is a different thing and is NOT the "Dynamic 3.0" GGUF which I am now…
-
I mostly use local models when the data has personal information. Earlier this year, I felt the coding quality was still not as good as Claude Code.One thing that works for me is to ask the local model to make some fake data with the same format, let Claude Code work on the fake data, and then bring the code back and run it locally on the real…
-
Are there benchmarks for the various Qwen3.8-27B quants that actually measure writing code, maybe even with multiple steps? Low KL divergence does not mean much when the model gets stuck in doom loops all the time.I could of course download and test myself, but that would take days with my internet connection.
-
"We also made some smaller UD-1bit quants with UD-IQ1_S being 6.2GB (without MTP) which retain around 72% top-1% accuracy yet being 89% smaller" This is crazy! But has anyone tried these lower quants on real projects?
-
Might be off-topic but: is it possible to perform such a quantization on Apple devices? Something like Mac Studio Ultra M1 (even if it would take weeks/months)?
-
Hey Unsloth, your gguf are the first ones I look for when I want to download a gguf model. Today I was trying in fact to see, what's the smallest Qwen3.8-27B that I could run and get good results, say restricting it to 16GB of ram.. so I went, pick up the Qwen3.8-27B-UD-IQ2_XXS.gguf and them BAM, error on MTP... now I understand why after reading…
-
H
Unsloth Dynamic 3.0 GGUFs L: https:// unsloth.ai/docs/basics/dynamic -3.0-ggufs C: https:// news.ycombinator.com/item?id=4 9365443 posted on 2026.08.19 at 14:36:45 (c=0, p=4)