Adaptation of GLM-5.3-Flash published as Jev-style decision model
9 Sep 26 11:49 AM · 2d ago · 1 post · 2 comments · 1 source · development 9 of 9
Another OSS variant emerges, this time adapting Alibaba's GLM-5.3-Flash into a System One decision model via PrivateMode's blog.
“I can't wait for a version of this to come along that supports images. I want to build a feature into Digital Carrot for creating AI goals where you can create a daily goal to, for example, "empty the dishwasher" that you would then verify by taking a picture of the empty dishwasher at the end of the day.”
newswangerd, Hacker News commenter · hn ↗TypeSafe Creator of Jev decision modelOllaya maintainers Open-source implementation leadsTinyJev maintainers Open-source implementation leadsValen authors Multimodal extension developers
The whole story postscomments the bright band is this development · numbered dots are the others · click one to jump
What was reported 1 claim about this development
-
first by HN Frontpage, 2d ago
What people said 17 voices · best of 18 · verbatim
-
I imagine it’s not so hard to optimize a model for this use case.Off the top of my head, I would skip all the modern linear attention / state space stuff and use classical attention. But run prefill in a fully sliding-window mode so that “state” tokens simply don’t attend to far away tokens, or maybe also allow everything to attend to the first…
-
I have run tests with qwen 3.8 and gemma 4 in a way similar to this post (based on an open source project that also does this with gemma4).Getting competitive accuracy with Jev is fairly easy, if by accuracy you mean that the highest weighted answer is the right one. GLM 5.3 is complete overkill, much smaller llms will doWhat Jev brings to the…
-
20k tok/sec prefill on B200/B300 isn't particularly noteworthy for medium-sized models like GLM-5.3-Flash, vLLM and SGLang achieve it on a reasonable number of models, especially at NVFP4.50k tok/sec is pretty impressive though.But... when you were doing your measurements, were you using the same random book excerpt? If you were potentially…
-
> I'm not sure what this means for AI startups if their innovations can be copied by OSS so quicklyThis particular "innovative" concept already has a rich, open research background. What Jev appears to have done is scale that up a bit and isolate good training data, which results in a great product but not really something impossible to imitate…
-
You can turn any sufficiently smart LLM into yes/no decision model or equivalent. I already have an existing workflow with a two paragraph detailed prompt, that sends pages of stuff to an LLM and asks it to return only 7 JSON objects. Several of those objects are binary "yes or no" choices of like, whether the content contains certain things.You…
-
It seems people are just guessing at the architecture behind Jev. Obviously the functionality itself is easy to replicate, but why Jev seems to be making such a splash (beyond the doh! factor of it's huge applicability) is the ultra-low cost and speed, which may be due to architecture.The Laya model compared in TFA shows one way Jev may be getting…
-
I can't wait for a version of this to come along that supports images. I want to build a feature into Digital Carrot for creating AI goals where you can create a daily goal to, for example, "empty the dishwasher" that you would then verify by taking a picture of the empty dishwasher at the end of the day. You can do this with current LLMs, but…
-
There is also a small wrapper over an unmodified llamacpp which implements something like this (raw label softmax, not calibrated!) on any llamacpp supported model. It is different than Leya/Kime here but can be useful for someone too.
-
was the answer correct?i have tested jev for my use cases and its horrendously wrong, but then the follow up from jev's team is "oh, you need to boil the question down further". it's a spiral of how much do you wanna dumb down the ask so that it answers it correctly. i'll pass for now.also, 30k input tokens is a lot.
-
Everyone is doing this to emulate Jev, but...I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.
-
Qwen3-Next-80B-A3B can already run on a 16GB M1 MacBook at around 3–5 tok/s using aggressive memory management. Could a Jev-style controller push that to 100 tok/s on an M1?
-
Is this a joke? “Jev-like” properties? People have been using LLMs as classifiers or rankers in a similar way for ages. I feel like we’re losing our minds
-
If you are using an autoregressive decoder (which glm is) it is not “jev-like”. You lose all of the speed advantages that Jev has.
-
I personally love Privatemode's approach. Having Jev-like speed for confidential ai use cases is a huge enabler.
-
Isnt this obvious ? I would have thought people would try such things before deciding they need something like Jev
-
how is Jev cheaper if I can run locally. 0.5% prefill, 0.1% decode, 99.4% cached, latency is <20ms
-
That's not true. I ran Cerebras as an experimental ultrafast Jev and it was faster.
All 9 developments of Open-source Jev clones proliferate weeks after TypeSafe's… →
Hacker NewsMastodonNewswiresLobsters