GLM announces proprietary inference infrastructure
1 Sep 17 · 10d ago · 1 article · 3 posts · 3 sources · development 1 of 1
GLM announced it has built its own inference infrastructure, moving away from reliance on third-party cloud providers for model serving.
GLM AI company
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
What was reported 1 claim about this development
-
first by HN Best, 10d ago · also HN Frontpage
What people said 23 voices · verbatim
-
I have a similar approach where I optimize kernels and find numerical differences between the CPU oracle and CUDA kernels using an automated AI agent in a feedback loop. Usually it solves numerical problems easily (it compares outputs of every layer and finds where they diverge), but so far no matter how many different SOTA models I throw at it…
-
Different angle on the same model: the full GLM-5.3 (744B MoE, 4-bit experts, 434 GB on disk) runs on a single MacBook Pro M5 Max with 128 GB by streaming the experts from NVMe SSDs instead of keeping them in memory.One drive gives about 2 tok/s; striped across four drives it reaches 3.5 tok/s with byte-identical output, and our best internal…
-
Nobody pretended that Chinese firms would just lay back and twiddle their thumbs when faced with import restrictions. The question was whether they would be far enough to be able to catch up without much issue or so far behind that they wouldn’t ever effectively catch up or that by the time they did, it wouldn’t matter.Half-arsed export…
-
US chip export winners and losers:Winners: Huawei, SMIC, CXMT,Chinese ASML-competitors, OpenAI, Anthropic, Amazon, Microsoft, Google, Meta.Losers: Chinese AI labs, Nvidia, AMD, TSMC, Micron, SK Hynix, Samsung, Intel.Any company that depends on Nvidia hardware such as OpenAI, Anthropic, AWS are winners. It means less competition for Nvidia chips…
-
This is a really funny sounding post. They sound like they just found out that increasing your automation gives you increased capabilities at faster speeds. They also sound like they just realized AI makes hard things easier.But what really kills me is the idea that these companies are using Python for production inference. I mean really? Have you…
-
Interesting that the tone of announcements between US and Chinese providers is converging.GLM has in the past been more technical rather than speculation about future development on RSI etc.Also curious whether those 100k accelerators are entirely locally made. If that's genuinely end to end on all components including lithography, memory, design…
-
Pretty impressive to see the amount of performance they can squeeze out of the same hardware. I suspect the same process will play out for all combinations of LLMs, inference providers and hardware stacks. This should bring down the cost of inference for the providers by an order of magnitude in the next year and lead to fantastic margins for…
-
It was evident that this will happen.> Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically…
-
China themselves recognize this. After Trump relaxed sanctions and allowed NVIDIA H200 sales to China on a case by case basis, the Chinese government stepped in to essentially block it!In addition to Huawei who make the Ascend series that Ziphu are using, there are also at least a half dozen or so other Chinese companies also making their own AI…
-
The first part of the article reads like Z.AI is trying to get their piece of the “national security concern” pie.The way these “AI is too powerful now” articles read about Mythos, Fable, GLM, etc is completely incongruent with my experience using them. It feels like they are all trying to position themselves to influence government policy.
-
If only this infrastructure could handle all the traffic. I've tried using glm via z.ai - and it's a snail kind of slow.And at the same time you have pretty strict limits to your usage, so in many cases you can't even let it work all night, as you will reach your limit faster than that.
-
Where the US sees itself penalizing China with an export restriction, China sees the US gifting it with zero-political-cost “protective” import tariff.You can’t really hurt a country that has a culture with a positive attitude toward growth.
-
"We implemented a series of aggressive memory optimizations, including..." This whole thing sounds like industrial scale auto-research, but done by people who actually know what they are doing.
-
"Necessity is the mother of invention".If someone is capable of doing something, and your goal is to prevent them for doing it, the worst thing you can do is to make it necessary for them to do it.
-
Given the huge amount of money being spent on AI chips in the US, what prevents US AI labs from doing the same level of software optimization? It could be a solve for some of the capacity constraints.
-
We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.
-
Yes, but the lag time is crucial. The NY Times estimated that China has 10% the compute capacity of the US. They may catch up, but by that time the US labs could have too much of an advantage
-
I might be missing something but when I went to their site they are more expensive than Claude. Why would I pick GLM over Claude? Is it they just offer more tokens in their plans?
-
Yup. That was really short sighted. And good for China. And actually the overall global market market since supply will augment and competition will decrease pricing as well.
-
US chip export restrictions may actually be an advantage for China's AI Infrastructure. Chinese companies are forced to speed up developing their own AI chips
-
Plot twist: the GLM optimization agent figured out that it can hack and use NVIDIA GPUs on a US Cloud provider and make the inference 10x faster.
-
The absolute goddamn hubris of the USA as a whole would be astonishing if it weren't so fucking stupid
-
Of course it is - how else could Huawei and a number smaller companies compete with NVIDIA.
All 1 developments of GLM builds proprietary inference infrastructure →
Hacker NewsMastodonNewswires