Fireworks releases Ember-1, distilled reasoning model at 40% fewer tokens
Fireworks AI released Ember-1, a specialized model based on Kimi K3 that maintains quality while cutting token use by 40% through efficient reasoning.
What to know
- Fireworks released Ember-1, achieving Kimi K3-level quality at 40% lower token cost through specialized training on efficient reasoning.
- The model addresses a real developer pain point: reasoning models waste tokens on unnecessary internal thought, compounding costs in multi-turn and agentic workflows.
- Community reaction focuses on tension between Fireworks' shift into proprietary research and Moonshot's more open approach to sharing model improvements.
- The release exemplifies a broader trend: developers are increasingly training small, specialized models for specific tasks rather than relying on general-purpose reasoning models.
The dispute Whether companies building on open-source models have an ethical or practical obligation to release their improvements, or whether proprietary enhancements are a legitimate business strategy. · positions read across 40 posts and comments
Fireworks' proprietary approach to Kimi K3 improvements contradicts open-source values that built their business.
-
“Moonshot has been openly improving, sharing research, and providing weights for the models that make up your entire bottom line, and the moment you can improve them in reciprocal it's closed weights, "this is our own proprietary" nonsense?”
nxtfari · Hacker News ↗
The trend toward specialized, cost-efficient models represents genuine progress and accessibility for developers.
-
“This is the golden age of model training… I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!”
GodelNumbering · Hacker News ↗
Questions remain about Fireworks' strategy and whether this represents a sustainable shift from infrastructure provider to model company.
-
“What am I missing here? I think of fireworks as an inference provider serving open weights model. The value that they primarily provide to customers is that (i) they improve reliability by balancing across a bunch of clouds/neoclouds, (ii)…”
tangled · Hacker News ↗
Fireworks Research Model developer and inference providerMoonshot (Kimi creator) Source model provider
How it unfolded 1 development · click the chart to see its coverage articlespostscomments
-
1
Commenters highlight broader trend in cost-efficient model specialization
Discussion emerged around the growing viability of training small, task-specific models instead of relying on large reasoning models. Commenters shared examples of successful custom training, including a model for Bash command translation and work on Qwen model specialization.
“This is the golden age of model training… I got a surprisingly good model for my task! The total active time I spent was a few hours.”
— GodelNumbering -
What am I missing here? I think of fireworks as an inference provider serving open weights model. The value that they primarily provide to customers is that (i) they improve reliability by balancing across a bunch of clouds/neoclouds, (ii) they get better pricing by buying capacity in bulk, and (iii) they reduce operational costs. So far so good.I…
2 more of the top 3 · 40 posts in this stretch
-
N
Ember-1: https:// fireworks.ai/blog/ember-1 Discussion: http:// news.ycombinator.com/item?id=4 9868830
-
I have been trying a mix of fine-tuning and I am amazed that most people do not see this coming.A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on…
-
-
background
Community questions Fireworks' move to proprietary models — Hacker News commenters raised concerns about Fireworks shifting from an open-weights inference provider to developing proprietary models. A commenter noted the tension with Moonshot's more open research ethos, while others questioned whether weights would be released.
-
background
Fireworks releases Ember-1 reasoning model — Fireworks AI announced Ember-1, a specialized model built on Kimi K3 that delivers equivalent quality with 40% fewer tokens. The model was trained to cut unnecessary reasoning while preserving critical thinking, tested across benchmarks and live production workloads.
Also covered reported alongside — the timeline has no entry for these yet
-
2 outlets Ember-1
first by HN Best, 1d ago · also HN Frontpage
What people are saying 21 voices from 1 site · best of 40 · verbatim
- Will Fireworks release the Ember-1 weights, or is this a closed proprietary model despite being trained on Kimi K3?
- Does Fireworks' shift toward proprietary research change its viability as a trusted provider for customer data and models?
- Yesterday
-
I vaguely recall a project from a while back that did something similar without LLMs.I’m really pushing my recall, but I want to say it was written in Ruby and stored pre-configured commands that it just did traditional search over.I vaguely recall it working okay because 99.99% of the questions people asked were the same (“tar command to gzip a…
-
I use this solution for your exact use case:I have a single command that fires up llama.cpp on cpu only using gemma4 e2b, answers a single question from the command line and exits. This takes about 3 seconds to load from an SSD, and is smart enough to solve exactly these "remind me of the syntax" scenarios if you dont wanna switch to a browser.
-
I don't know, I have a weird feeling that Ember-1 is a distil/combination of Qwen + Gemini 3 Flash.Looking at the hamsters drawings in this comparison, never saw any other model make such a similar version:
- Sep 27
-
Here's a similar project for those who want to replicate: https://github.com/ThorOdinson246/whatisit-nl2shNot my project
-
I work at Fireworks and it's cool to see this was posted.I'd be interested to hear what people found most interesting about Ember, and what kinds of follow-up research or educational material would be useful to you all?
-
I don't see why I would be interested in this model, considering the price difference. They advertise that it's the same as Kimi K3 in half the tokens. But the pricing is double the pricing of K3. So why do I care if it uses fewer tokens, if I'm paying double per token?
-
Seriously, I'm using a Qwen 3.8 27B on the homelab, distilled from supposed Fable traces. Regardless, the difference is notable, less thinking, better output. Distilled / heavy quant is better than the original (imv)https://huggingface.co/vwdubb/Qwen3.8-27B-Fable-Distill-NVFP...side quest, are fable distillations only wrong when it's another…
-
This is undoubtedly great. But most of the inference cost today for dominant use cases (agentic coding) are in the prefill, not the decode. This is one of the reasons that DeepSeek is so aggressively optimizing prefill and caching.
-
Given what an experience i had with Ember-2… I’m not sure I’d want to engage with its predecessor.
-
The more I learn about Fireworks the more unsavory they seem as a company. I don’t care what the license says, Moonshot has been openly improving, sharing research, and providing weights for the models that make up your entire bottom line, and the moment you can improve them in reciprocal it’s closed weights, “this is our own proprietary”…
-
It’s the first time I know fireworks has a team doing model research. I do have a complex mood in that. On one hand, I’m always happy to see improvement of OSS models, whether that’s on intelligence or cost-efficiency. On the other hand, I would be a little worried about using fireworks as my API provider. Till the moment I saw this news, I had…
-
Been thinking about the feasibility of training a model using synthetic thinking traces that were reduced to caveman-speak prior to being used for training. Seems like it would be fairly easy to generate plenty of suitably lobotomized synthetic traces with a pair of cheap-ish models. Or even just using good old fashioned NLP to aggressively remove…
-
> The problem: thinking models think too muchI see that with Opus 5, it started thinking like crazy in the last few days , I don't think my workflow is that complicated, still it gets into thinking mode and stays there
-
That's a really impressive result. There are all kinds of small tasks like this I use an LLM for, but theoretically if you broke all the sub-use cases into local-only models, and had something lightweight that routed to the right model, you could have faster and cheaper workflows. E.g. something trained on the linux man pages for common commands…
-
Curious about how you generated the training data? Was it just asking an existing model to generate a bunch of examples?I ask cause would this be a kind of model distillation?I have a small model I'm looking to train on some data, and I have some real live data but I'd love to be able to extend it.
-
Over on /r/LocalLLaMA there's a group that's been getting popular doing the same thing for the Qwen 27B (and other) models. -
-
This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it…
-
> The problem: thinking models think too muchThis is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinkingIt’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy…
-
Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs. Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds) Sol is at 2/10 vs kimi’s 3/15
-
Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?
-
Well done, and great iteration.The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).