Researchers contest the benchmark claim with contrary test results
2 Sep 16 11:15 AM · 11d ago · 1 post · 1 comment · 2 sources · development 2 of 5
Hacker News commenter TuxSH reports benchmarking DeepSeek v4.1 Flash against GLM 5.3 on a Nintendo 3DS kernel decomposition task. GLM 5.3 found almost all vulnerabilities in 30 minutes for $22, while DeepSeek found only one vulnerability in 40 minutes for $2, suggesting the model may perform better on low-hanging fruit targets.
“GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min. Perhaps DS works better where targets have low-hanging fruits than can be found fast?”
TuxSHEnclave.ai Security research firmDeepSeek AI model provider
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 1 voice · verbatim
-
I find this - or perhaps the title - a bit surprising.I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln…
All 5 developments of DeepSeek v4.1 Flash Achieves Perfect Score on AI Hacking… →
Hacker NewsMastodonNewswires