AMD acquires Taalas to embed AI models directly in silicon
The House of Zen's acquisition of the Toronto-based chip startup aims to deliver inference speeds up to 17,000 tokens per second by baking model weights into silicon.
Conversation activity · last 3 days peak 3/hr
Summary, timeline and people extracted by Claude from 11 items across 3 sources · 4h ago. Quotes are verbatim.
AMD announced the acquisition of Taalas, a Toronto-based AI chip startup that etches model weights directly into silicon to dramatically accelerate AI inference. The company's model-specific integrated circuits (MSICs) have demonstrated throughput of 16,960 tokens per second on Meta's Llama 3.1 8B model in early benchmarks. The deal positions AMD to compete with Nvidia's dominance in AI hardware and follows the pattern of Nvidia's $20 billion licensing deal with Groq.
- AMD acquired Taalas, a Toronto startup that bakes AI model weights directly into silicon, achieving inference speeds of 16,960 tokens per second on Llama 3.1 8B—far exceeding current GPU and accelerator benchmarks.
- The technology uses model-specific integrated circuits (MSICs) with mask-ROM fabric for weights and SRAM for KV caches, enabling efficiency gains that could require just 50 accelerators for trillion-parameter models compared to thousands of GPUs.
- The deal positions AMD to directly challenge Nvidia's AI hardware dominance and echoes Nvidia's $20 billion licensing deal with Groq, signaling a shift toward specialized, purpose-built inference silicon.
- Community discussion highlights potential transformative impact on robotics, IoT, edge deployment, and emerging use cases enabled by 100x+ speed improvements, though raises questions about model obsolescence and the rapidly evolving AI landscape.
How it unfolded
-
Analysis Discussion of new use cases enabled by speed
Commenters highlight that faster inference opens new classes of user experiences and applications, similar to how faster internet enabled SaaS and streaming rather than just faster web browsing.
“When technology gets faster, it opens up whole new classes of UX that were hard to predict. For example, faster internet didn't mean being able to view 100x as many HTML4 web pages. It brought SaaS, streaming media and interactivity.”
dave1010uk · Hacker News ↗ -
Reaction Community discusses future implications
Commenters speculate about the long-term impact of dramatically faster inference, including questions about model obsolescence and the pace of AI advancement.
“Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.”
linzhangrun · Hacker News ↗ -
Reaction Tech community reacts on Hacker News
Hacker News commenters begin discussing the implications of the acquisition, noting its significance for robotics, IoT, and edge AI applications.
“People are not talking enough how huge this is for robotics and IoT. Current robotics arhitectures are limited by tok/sec. How cares if its not a Fable model? This move undercuts NVIDIA directly.”
trash_cat · Hacker News ↗ -
Event AMD announces Taalas acquisition
AMD announces the acquisition of Taalas at market close on Thursday. The company intends to pair Taalas-based chips with its Instinct-based Helios racks for disaggregated AI workload processing.
“AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload.”
Vamsi Boppana, AMD SVP of AI · Techmeme ↗ - 27 weeks quiet
-
Report Taalas HC1 chip revealed
Taalas unveils its first test chip, the HC1, fabricated on TSMC's 6nm process, achieving 16,960 tokens per second on Llama 3.1 8B—48x faster than Nvidia GPUs and 8.5x faster than Cerebras accelerators at the time.
- 161 weeks quiet
-
Event Taalas founded
Taalas is founded in Toronto as an AI chip company focused on a novel inference approach.
What people are saying verbatim
“AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload.”
Vamsi Boppana, AMD SVP of AI · The Register ↗ · Aug 5
“People are not talking enough how huge this is for robotics and IoT. Current robotics arhitectures are limited by tok/sec. This move undercuts NVIDIA directly.”
trash_cat, Hacker News commenter · Hacker News ↗ · Aug 6
“I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition.”
LarsDu88, Hacker News commenter · Hacker News ↗ · Aug 5
“Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.”
linzhangrun, Hacker News commenter · Hacker News ↗ · Aug 6
“When technology gets faster, it opens up whole new classes of UX that were hard to predict. For example, faster internet didn't mean being able to view 100x as many HTML4 web pages. It brought SaaS, streaming media and interactivity.”
dave1010uk, Hacker News commenter · Hacker News ↗ · Aug 6
“I see a lot of discourse about it being fast-to-deprecation. But I see it a different way personally. Modern LLMs are trying to do more with less. The significantly faster speed means quicker iterations.”
dabbz, Hacker News commenter · Hacker News ↗ · Aug 6
Voices from the web unedited
-
I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device."Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power…
-
I'm surprised there's not more discussion about potential inflection points here. When technology gets faster, it opens up whole new classes of UX that were hard to predictFor example, faster internet didn't mean being able to view 100x as many HTML4 web pages. It brought SaaS, streaming media and interactivity.I'm not good at predicting, but some…
-
I see a lot of discourse about it being fast-to-deprecation. But I see it a different way personally.Modern LLMs are trying to do more with less. Focus on doing the right thing the first time. Even if we squeeze dumb LLMs, the significantly faster speed means quicker iterations. So a bad decision doesn't cost the time and inference costs that it…
-
I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition.Baking models onto silicon would've been the next logical move to get a moat.Google is already doing this and has an experimental project on top of already having TPUs and cramming their…
-
People are not talking enough how huge this is for robotics and IoT. Current robotics arhitectures are limited by tok/sec. How cares if its not a Fable model?This move undercuts NVIDIA directly.
-
What I like about this, is that it significantly increases the probability of a sci-fi scenario where you're picking up a hot chip on the black market; rumor has it, Mythos 9 weights baked in...
-
Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.