conv.

All stories
AIActive · 37h

Fireworks releases Ember-1, distilled reasoning model at 40% fewer tokens

Fireworks AI released Ember-1, a specialized model based on Kimi K3 that maintains quality while cutting token use by 40% through efficient reasoning.

What to know

  • Fireworks released Ember-1, achieving Kimi K3-level quality at 40% lower token cost through specialized training on efficient reasoning.
  • The model addresses a real developer pain point: reasoning models waste tokens on unnecessary internal thought, compounding costs in multi-turn and agentic workflows.
  • Community reaction focuses on tension between Fireworks' shift into proprietary research and Moonshot's more open approach to sharing model improvements.
  • The release exemplifies a broader trend: developers are increasingly training small, specialized models for specific tasks rather than relying on general-purpose reasoning models.

The dispute Whether companies building on open-source models have an ethical or practical obligation to release their improvements, or whether proprietary enhancements are a legitimate business strategy. · positions read across 40 posts and comments

many voices

Fireworks' proprietary approach to Kimi K3 improvements contradicts open-source values that built their business.

  • “Moonshot has been openly improving, sharing research, and providing weights for the models that make up your entire bottom line, and the moment you can improve them in reciprocal it's closed weights, "this is our own proprietary" nonsense?”

    nxtfari · Hacker News ↗
many voices

The trend toward specialized, cost-efficient models represents genuine progress and accessibility for developers.

  • “This is the golden age of model training… I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!”

    GodelNumbering · Hacker News ↗
some voices

Questions remain about Fireworks' strategy and whether this represents a sustainable shift from infrastructure provider to model company.

  • “What am I missing here? I think of fireworks as an inference provider serving open weights model. The value that they primarily provide to customers is that (i) they improve reliability by balancing across a bunch of clouds/neoclouds, (ii)…”

    tangled · Hacker News ↗

Fireworks Research Model developer and inference providerMoonshot (Kimi creator) Source model provider

How it unfolded 1 development · click the chart to see its coverage articlespostscomments

Peak 8 pieces in one hour at Sep 27, 4 PM; 43 pieces over 38 hours (2 articles · 3 posts · 38 comments) Sep 27, 1 PM — 3 pieces · 2 articles · 1 post — Newswires 2, Hacker News 1Sep 27, 2 PM — 7 pieces · 7 comments — Hacker News 7Sep 27, 3 PM — 5 pieces · 5 comments — Hacker News 5Sep 27, 4 PM — 8 pieces · 1 post · 7 comments — Hacker News 7, Mastodon 1Sep 27, 5 PM — 3 pieces · 3 comments — Hacker News 3Sep 27, 6 PM — 1 piece · 1 comment — Hacker News 1Sep 27, 7 PM — 3 pieces · 3 comments — Hacker News 3Sep 27, 8 PM — 1 piece · 1 comment — Hacker News 1Sep 27, 9 PM — 1 piece · 1 comment — Hacker News 1Sep 27, 10 PM — 1 piece · 1 comment — Hacker News 1Sep 27, 11 PM — quietYesterday, 12 AM — quietYesterday, 1 AM — quietYesterday, 2 AM — quietYesterday, 3 AM — quietYesterday, 4 AM — 3 pieces · 3 comments — Hacker News 3Yesterday, 5 AM — quietYesterday, 6 AM — 2 pieces · 2 comments — Hacker News 2Yesterday, 7 AM — 1 piece · 1 comment — Hacker News 1Yesterday, 8 AM — 1 piece · 1 post — Mastodon 1Yesterday, 9 AM — quietYesterday, 10 AM — 2 pieces · 2 comments — Hacker News 2Yesterday, 11 AM — 1 piece · 1 comment — Hacker News 1Yesterday, 12 PM — quietYesterday, 1 PM — quietYesterday, 2 PM — quietYesterday, 3 PM — quietYesterday, 4 PM — quietYesterday, 5 PM — quietYesterday, 6 PM — quietYesterday, 7 PM — quietYesterday, 8 PM — quietYesterday, 9 PM — quietYesterday, 10 PM — quietYesterday, 11 PM — quietToday, 12 AM — quietToday, 1 AM — quietToday, 2 AM — quiet 1
4 PMyesterday8 AM4 PMnow · 3:09 AM ET
  1. 1

    Commenters highlight broader trend in cost-efficient model specialization

    Discussion emerged around the growing viability of training small, task-specific models instead of relying on large reasoning models. Commenters shared examples of successful custom training, including a model for Bash command translation and work on Qwen model specialization.

    “This is the golden age of model training… I got a surprisingly good model for my task! The total active time I spent was a few hours.”
    — GodelNumbering
    • What am I missing here? I think of fireworks as an inference provider serving open weights model. The value that they primarily provide to customers is that (i) they improve reliability by balancing across a bunch of clouds/neoclouds, (ii) they get better pricing by buying capacity in bulk, and (iii) they reduce operational costs. So far so good.I…

      tangledHacker News1d agoview on Hacker News ↗
    2 more of the top 3 · 40 posts in this stretch
    • newsyc500@mastodon.social

      Ember-1: https:// fireworks.ai/blog/ember-1 Discussion: http:// news.ycombinator.com/item?id=4 9868830

      newsyc500@mastodon.socialMastodon18h agoview on Mastodon ↗
    • I have been trying a mix of fine-tuning and I am amazed that most people do not see this coming.A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on…

      brainlessHacker News1d agoview on Hacker News ↗
    all of them →
  2. background

    Community questions Fireworks' move to proprietary models — Hacker News commenters raised concerns about Fireworks shifting from an open-weights inference provider to developing proprietary models. A commenter noted the tension with Moonshot's more open research ethos, while others questioned whether weights would be released.

  3. background

    Fireworks releases Ember-1 reasoning model — Fireworks AI announced Ember-1, a specialized model built on Kimi K3 that delivers equivalent quality with 40% fewer tokens. The model was trained to cut unnecessary reasoning while preserving critical thinking, tested across benchmarks and live production workloads.

Also covered reported alongside — the timeline has no entry for these yet

  1. 2 outlets Ember-1

    first by HN Best, 1d ago · also HN Frontpage

What people are saying 21 voices from 1 site · best of 40 · verbatim

Still unanswered
  • Will Fireworks release the Ember-1 weights, or is this a closed proprietary model despite being trained on Kimi K3?
  • Does Fireworks' shift toward proprietary research change its viability as a trusted provider for customer data and models?