Commenters highlight broader trend in cost-efficient model specialization
1 Sep 27 · 1d ago · 2 articles · 2 posts · 18 comments · 3 sources · development 1 of 1
Discussion emerged around the growing viability of training small, task-specific models instead of relying on large reasoning models. Commenters shared examples of successful custom training, including a model for Bash command translation and work on Qwen model specialization.
“This is the golden age of model training… I got a surprisingly good model for my task! The total active time I spent was a few hours.”
GodelNumberingFireworks Research Model developer and inference providerMoonshot (Kimi creator) Source model provider
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
Reported in the same hours no headline names this development itself — these 1 claim were published in its stretch
-
2 outlets Ember-1
first by HN Best, 1d ago · also HN Frontpage
What people said 24 voices · best of 40 · verbatim
-
What am I missing here? I think of fireworks as an inference provider serving open weights model. The value that they primarily provide to customers is that (i) they improve reliability by balancing across a bunch of clouds/neoclouds, (ii) they get better pricing by buying capacity in bulk, and (iii) they reduce operational costs. So far so good.I…
-
N
Ember-1: https:// fireworks.ai/blog/ember-1 Discussion: http:// news.ycombinator.com/item?id=4 9868830
-
I have been trying a mix of fine-tuning and I am amazed that most people do not see this coming.A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on…
-
It’s the first time I know fireworks has a team doing model research. I do have a complex mood in that. On one hand, I’m always happy to see improvement of OSS models, whether that’s on intelligence or cost-efficiency. On the other hand, I would be a little worried about using fireworks as my API provider. Till the moment I saw this news, I had…
-
> The problem: thinking models think too muchThis is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinkingIt’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy…
-
This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it…
-
Seriously, I'm using a Qwen 3.8 27B on the homelab, distilled from supposed Fable traces. Regardless, the difference is notable, less thinking, better output. Distilled / heavy quant is better than the original (imv)https://huggingface.co/vwdubb/Qwen3.8-27B-Fable-Distill-NVFP...side quest, are fable distillations only wrong when it's another…
-
I don't know, I have a weird feeling that Ember-1 is a distil/combination of Qwen + Gemini 3 Flash.Looking at the hamsters drawings in this comparison, never saw any other model make such a similar version:
-
That's a really impressive result. There are all kinds of small tasks like this I use an LLM for, but theoretically if you broke all the sub-use cases into local-only models, and had something lightweight that routed to the right model, you could have faster and cheaper workflows. E.g. something trained on the linux man pages for common commands…
-
I vaguely recall a project from a while back that did something similar without LLMs.I’m really pushing my recall, but I want to say it was written in Ruby and stored pre-configured commands that it just did traditional search over.I vaguely recall it working okay because 99.99% of the questions people asked were the same (“tar command to gzip a…
-
The more I learn about Fireworks the more unsavory they seem as a company. I don’t care what the license says, Moonshot has been openly improving, sharing research, and providing weights for the models that make up your entire bottom line, and the moment you can improve them in reciprocal it’s closed weights, “this is our own proprietary”…
-
Been thinking about the feasibility of training a model using synthetic thinking traces that were reduced to caveman-speak prior to being used for training. Seems like it would be fairly easy to generate plenty of suitably lobotomized synthetic traces with a pair of cheap-ish models. Or even just using good old fashioned NLP to aggressively remove…
-
I use this solution for your exact use case:I have a single command that fires up llama.cpp on cpu only using gemma4 e2b, answers a single question from the command line and exits. This takes about 3 seconds to load from an SSD, and is smart enough to solve exactly these "remind me of the syntax" scenarios if you dont wanna switch to a browser.
-
Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs. Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds) Sol is at 2/10 vs kimi’s 3/15
-
Curious about how you generated the training data? Was it just asking an existing model to generate a bunch of examples?I ask cause would this be a kind of model distillation?I have a small model I'm looking to train on some data, and I have some real live data but I'd love to be able to extend it.
-
Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?
-
I don't see why I would be interested in this model, considering the price difference. They advertise that it's the same as Kimi K3 in half the tokens. But the pricing is double the pricing of K3. So why do I care if it uses fewer tokens, if I'm paying double per token?
-
Well done, and great iteration.The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).
-
Given what an experience i had with Ember-2… I’m not sure I’d want to engage with its predecessor.
-
Over on /r/LocalLLaMA there's a group that's been getting popular doing the same thing for the Qwen 27B (and other) models. -
-
Here's a similar project for those who want to replicate: https://github.com/ThorOdinson246/whatisit-nl2shNot my project
-
I work at Fireworks and it's cool to see this was posted.I'd be interested to hear what people found most interesting about Ember, and what kinds of follow-up research or educational material would be useful to you all?
-
This is undoubtedly great. But most of the inference cost today for dominant use cases (agentic coding) are in the prefill, not the decode. This is one of the reasons that DeepSeek is so aggressively optimizing prefill and caching.
-
> The problem: thinking models think too muchI see that with Opus 5, it started thinking like crazy in the last few days , I don't think my workflow is that complicated, still it gets into thinking mode and stays there
All 1 developments of Fireworks releases Ember-1, distilled reasoning model at… →
Hacker NewsNewswiresMastodon