gertlabs provides internal evaluation positioning v4.1 Flash on performance spectrum
5 Sep 16 1:29 PM · 11d ago · 4 comments · 1 source · development 5 of 5
A commenter identifying as gertlabs shares internal evaluation results placing DeepSeek v4.1 Flash near Gemini 3.7 Flash on the cost-performance frontier when accounting for actual reasoning overhead, though it outperforms most open-weight models in agentic coding at lower cost.
“Because it's a heavy reasoner, it sits near Gemini 3.7 Flash on the Pareto front (not as cheap as the price suggests in practice).”
gertlabsEnclave.ai Security research firmDeepSeek AI model provider
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 4 voices · verbatim
-
We ran v4.1 Flash through our evaluations and found it to be smarter and faster than V4 Flash, with a commensurate price bump. Some notes:- Because it's a heavy reasoner, it sits near Gemini 3.7 Flash on the Pareto front (not as cheap as the price suggests in practice).- Closer than expected to the top open weights models (GLM 5.3 and Kimi K3) in…
-
I'm a deepseek fanboy, but has anyone else found flash to not meet expectations? I've found it to be wildly bad at doing as asked, overengineering, and always assuming instead of reading even if told to read things in full before doing anything. It makes wildly silly mistakes in code and so far has been quite frustrating to work with. Maybe its…
-
How do I make deepseek "hack" my source code? do I just start my coding agent in my directory and command it to "find vulnerabilities", or is there some more sophisticated software to do that?
-
we recently got this running in 192gb of vRAM and using it with the Klaudia harness has been incredible for driving out work that we'd need to use Opus and Fable for previously
All 5 developments of DeepSeek v4.1 Flash Achieves Perfect Score on AI Hacking… →
Hacker NewsMastodonNewswires