conv.

All stories
AIQuiet 40d · day 41

Unsloth releases Dynamic 3.0 quantization with >10% accuracy gain

New GGUF quantization method delivers better model quality at same file size for Qwen 3.8-27B.

What to know

  • Unsloth's Dynamic 3.0 quantization delivers >10% better top-1% accuracy than competitors at identical file sizes using improved calibration data and post-training techniques.
  • The smallest variant (UD-IQ1_S) achieves 72% accuracy at 6.2GB—89% smaller than full precision—but real-world coding performance remains unvalidated by users.
  • Community feedback highlights practical concerns: file versioning confusion, MTP removal trade-offs, and demand for code-generation benchmarks beyond KL divergence metrics.

The dispute Whether KL Divergence and top-1% accuracy metrics reliably predict practical code generation performance in production settings. · positions read across 16 posts and comments

many voices

Published metrics are insufficient without real-world code generation validation.

  • “Low KL divergence does not mean much when the model gets stuck in doom loops all the time.”

    johndough · Hacker News
some voices

Unsloth quants are high-quality and worth using with proper tooling.

  • “Your gguf are the first ones I look for when I want to download a gguf model.”

    xlayn · Hacker News
many voices

File management and metadata practices need improvement for practical deployment.

  • “It would be nice if unsloth published GGUFs would use a version number or something, because now I have multiple different files on local storage that otherwise have exactly the same name.”

    walrus01 · Hacker News

Unsloth Quantization library developer

The record 2 articles and posts · last 30 days

  1. summary covers to here · Aug 19, 10:48 PM · 1 piece above arrived after
  2. Unsloth Dynamic 3.0 GGUFs press · HN Best · jonesy827 · 40d ago · +1 outlet
  3. Unsloth Dynamic 3.0 GGUFs: https:// unsloth.ai/docs/basics/dynamic -3.0-ggufs Discussion: http://… post · Mastodon · newsyc250 · 40d ago

What people are saying 16 voices from 2 sites · verbatim

Still unanswered
  • How do the smaller quants (UD-IQ1_S, UD-IQ2_XXS) perform on real code generation tasks with multiple reasoning steps?
  • What is the performance impact of removing MTP from smaller quantizations on inference speed for memory-constrained devices?