
AMD buys Taalas, the startup that hardwires AI models into silicon
AMD is acquiring Taalas, a Toronto startup founded in 2023 that fixes a single model's weights into a chip's metal layers rather than streaming them from memory — its first part runs Meta's Llama 3.1 8B at a claimed 17,000 tokens per second per user. Terms were undisclosed; Taalas had raised $219M since founding. Seven months after Nvidia paid $20B for Groq's assets, inference silicon is splitting into two tiers — flexible GPUs for new models, fixed-function chips for the few models running at scale — and AMD is folding Taalas into its Instinct and Helios rack roadmap.
Source: cnbc.com ↗
There's no one-size-fits-all as it comes to chips.
Lisa Su, AMD CEO
Why this matters
- → Inference silicon diverging into specialized fixed-function and flexible GPU tiers.
- → AMD racing Nvidia to lock in models at scale via hardware.
- → Custom silicon promises 17k tokens/sec — vastly faster for specific workloads.
The inference fork