conv.

All stories
AIQuiet 9d · day 11

GLM builds proprietary inference infrastructure

Chinese AI company GLM has developed its own inference infrastructure rather than relying on third-party cloud providers.

What to know

  • GLM has developed in-house inference infrastructure rather than depending on external cloud providers.
  • The move represents vertical integration in the AI industry as companies seek control over deployment and serving of their models.

GLM AI company

GLM builds proprietary inference infrastructure
z.ai

How it unfolded 1 development · click the chart to see its coverage articlesposts

Peak 11 pieces in 3h at Sep 17, 6 AM; 29 pieces over 11 days (2 articles · 4 posts · 23 comments) Sep 17, 3 AM — 5 pieces · 2 articles · 2 posts · 1 comment — Newswires 2, Hacker News 2, Mastodon 1Sep 17, 6 AM — 11 pieces · 1 post · 10 comments — Hacker News 10, Mastodon 1Sep 17, 9 AM — 8 pieces · 1 post · 7 comments — Hacker News 7, Mastodon 1Sep 17, 12 PM — 3 pieces · 3 comments — Hacker News 3Sep 17, 3 PM — 1 piece · 1 comment — Hacker News 1Sep 17, 6 PM — quietSep 17, 9 PM — quietSep 18, 12 AM — quietSep 18, 3 AM — quietSep 18, 6 AM — 1 piece · 1 comment — Hacker News 1Sep 18, 9 AM — quietSep 18, 12 PM — quietSep 18, 3 PM — quietSep 18, 6 PM — quietSep 18, 9 PM — quietSep 19, 12 AM — quietSep 19, 3 AM — quietSep 19, 6 AM — quietSep 19, 9 AM — quietSep 19, 12 PM — quietSep 19, 3 PM — quietSep 19, 6 PM — quietSep 19, 9 PM — quietSep 20, 12 AM — quietSep 20, 3 AM — quietSep 20, 6 AM — quietSep 20, 9 AM — quietSep 20, 12 PM — quietSep 20, 3 PM — quietSep 20, 6 PM — quietSep 20, 9 PM — quietSep 21, 12 AM — quietSep 21, 3 AM — quietSep 21, 6 AM — quietSep 21, 9 AM — quietSep 21, 12 PM — quietSep 21, 3 PM — quietSep 21, 6 PM — quietSep 21, 9 PM — quietSep 22, 12 AM — quietSep 22, 3 AM — quietSep 22, 6 AM — quietSep 22, 9 AM — quietSep 22, 12 PM — quietSep 22, 3 PM — quietSep 22, 6 PM — quietSep 22, 9 PM — quietSep 23, 12 AM — quietSep 23, 3 AM — quietSep 23, 6 AM — quietSep 23, 9 AM — quietSep 23, 12 PM — quietSep 23, 3 PM — quietSep 23, 6 PM — quietSep 23, 9 PM — quietSep 24, 12 AM — quietSep 24, 3 AM — quietSep 24, 6 AM — quietSep 24, 9 AM — quietSep 24, 12 PM — quietSep 24, 3 PM — quietSep 24, 6 PM — quietSep 24, 9 PM — quietSep 25, 12 AM — quietSep 25, 3 AM — quietSep 25, 6 AM — quietSep 25, 9 AM — quietSep 25, 12 PM — quietSep 25, 3 PM — quietSep 25, 6 PM — quietSep 25, 9 PM — quietYesterday, 12 AM — quietYesterday, 3 AM — quietYesterday, 6 AM — quietYesterday, 9 AM — quietYesterday, 12 PM — quietYesterday, 3 PM — quietYesterday, 6 PM — quietYesterday, 9 PM — quietToday, 12 AM — quietToday, 3 AM — quietToday, 6 AM — quietToday, 9 AM — quietToday, 12 PM — quietToday, 3 PM — quiet 1
Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25yesterdaynow · 5:40 PM ET
  1. 1
    1. first by HN Best, 10d ago · also HN Frontpage

    • I have a similar approach where I optimize kernels and find numerical differences between the CPU oracle and CUDA kernels using an automated AI agent in a feedback loop. Usually it solves numerical problems easily (it compares outputs of every layer and finds where they diverge), but so far no matter how many different SOTA models I throw at it…

      kgeistHacker News10d agoview on Hacker News ↗
    2 more of the top 3 · 23 posts in this stretch
    • Different angle on the same model: the full GLM-5.3 (744B MoE, 4-bit experts, 434 GB on disk) runs on a single MacBook Pro M5 Max with 128 GB by streaming the experts from NVMe SSDs instead of keeping them in memory.One drive gives about 2 tok/s; striped across four drives it reaches 3.5 tok/s with byte-identical output, and our best internal…

      ArgonautlabsHacker News10d agoview on Hacker News ↗
    • Nobody pretended that Chinese firms would just lay back and twiddle their thumbs when faced with import restrictions. The question was whether they would be far enough to be able to catch up without much issue or so far behind that they wouldn’t ever effectively catch up or that by the time they did, it wouldn’t matter.Half-arsed export…

      christina97Hacker News10d agoview on Hacker News ↗
    all of them →

What people are saying 20 voices from 1 site · best of 23 · verbatim