Community explores model deployment in browsers and on local hardware
2Sep 17 5:56 PM · 8d ago · 1 comment · 1 source · development 2 of 4
Developers report successfully running the model in browsers via WebGPU kernels and on consumer hardware including Mac Mini M2 and NVIDIA GPUs, with observed throughputs of 7-8 tokens/second on M2 and 143 tokens/second on RTX 5090.
prismml.com
“These are small enough that you can run them entirely in the browser”
Aurornis
PrismMLModel developer and publisherAurornisCommunity tester
The whole story articlespostscommentsthe bright band is this development · numbered dots are the others · click one to jump
These are small enough that you can run them entirely in the browser https://huggingface.co/spaces/webml-community/ternary-bonsai...Remember to clear the downloaded weights afterward.Like the last model, it's amazing they work as well as they do. Use it for any longer task and they fall apart spectacularly and in interesting ways.