Nandakishor M releases RL Agent, an open-source alternative
1 Sep 19 9:41 AM · 7d ago · 3 posts · 2 sources · development 1 of 1
Rather than remain bitter, the researcher built RL Agent, a completely open-source System 1 decision model using a bidirectional encoder. He claims it runs in 33–38 milliseconds on GPU, approximately 4x faster than Jev's 150 ms latency.
“And because we built it properly on a bidirectional encoder, our model runs in 33 to 38 milliseconds on a GPU, making it roughly 4x faster than Jev's published 150 ms latency, and it is 100% open-source.”
Nandakishor MNandakishor M Independent AI researcherDiogo Almeida TypeSafe AI founder, ChatGPT co-inventorTypeSafe AI Commercial AI startup
The whole story posts the bright band is this development · numbered dots are the others · click one to jump
What people said 7 voices · verbatim
-
> In my March 2025 work, I used frozen sequence embeddings combined with a separate PPO value network. It worked for turn-by-turn sales prediction, but it was not end-to-end **and could not handle dynamic new questions at runtime**. (emphasis mine) Maybe this is just my lack of ML knowledge showing, but isn’t this kind of the entire selling point…
-
B
🔬 The best bytes of # science & # tech across the # fediverse “I Built Non-Autoregressive Decision Models a Year Ago․ Then a Frontier Lab Called It a 'Breakthrough' dev․to/nandakishor_m_6cc0adfde9f/i-built-non-autoregressive-decision-models-a-year-ago-then-a-frontier-lab-called-it…” https:// dev.to/nandakishor_m_6cc0adfde…
-
I can be wrong, but my impression is that they are building a generic classifier, for which you supply the task in a natural language and the schema of the output, and get some (hopefully, good) predictions. Classifiers are obviously not a new thing, but the old way is to gather a labeled dataset specifically for your task and then train a model…
-
Should be merged with https://lobste.rs/s/ojukrw/laya_33ms_multilingual_system_1_decision, I think
-
I haven’t been in the trade for 20 years but, in my days, I solved the "decision problem" with simple bayesian learning. When the input is text to classify and the number of classification is low, bayesian are really really good.
-
It's both a pro and a con for Jev. It contains lots of world knowledge so you can go more high-level. On the other hand you can't tune it to your needs too much. If you can't sufficiently describe the rules you need and potentially override existing ideas, you won't benefit from Jev.
-
Their original approach had that limitation. It looks like the new one (continuing with lessons learnt from 2025) supports any questions.
All 1 developments of Researcher claims TypeSafe AI's Jev repackages his year-old… →
LobstersMastodon