Skepticism over "DeepSeek V4 Flash grade" claim amid demo failures
5Sep 18 3:38 PM · 7d ago · 1 post · 4 comments · 2 sources · development 5 of 7
As more users tested the demo and encountered routing errors, skepticism grew about the claimed performance parity with DeepSeek V4 Flash. One commenter noted the claim "seems far fetched" given observed failures.
“The "DeepSeek 4 Flash grade" claim seems far fetched.”
I tried this today for labelling - and for that task it was very bad MNLI was better - so you are going to need to match the use case for this pretty exactly. (at 29MB params one would expect that!) I'm obviously not saying labelling is a good use case :-) just adding a data point.Jev has put the cat amongst the pigeons so suddenly everyone is…
Pretty much matches my experience.> 'sleepy time' means sleeping → start_vacuum with room 'bedroom' to start cleaningThe "DeepSeek 4 Flash grade" claim seems far fetched.
What is 8 dash 29 MB? And the copy on the landing page is clearly AI generated with the “each layer a model of its own” stuff, makes little sense. The more I see AI generated copy the less it makes sense.
Really interesting project. The intelligence laddering and on-device tool calling are especially cool. Nice work getting this running across so many platforms!