⚡ Uncle Cat AI Radar
IndustryModels

AMD buys Taalas to etch AI models directly into silicon

The Toronto startup builds chips whose metal layers are finalised around one model's weights, promising inference without HBM, advanced packaging or liquid cooling.

AMD said on Thursday it has agreed to acquire Taalas, a Toronto chip startup founded in 2023 that builds processors around a single fixed model rather than around general-purpose compute. Terms were not disclosed. The deal is subject to regulatory approval and is expected to close in the fourth quarter.

What Taalas makes

Taalas calls its products "Hardcore Models." Once a model's weights are frozen, the company finalises a small number of a chip's metal layers to encode them, collapsing the usual separation between memory and compute onto one die. The claimed consequence is a serving system that needs no high-bandwidth memory, no advanced packaging, no 3D stacking and no liquid cooling — the four cost centres that dominate the bill of materials for a modern inference rack. Taalas says a previously unseen model can be turned into hardware in about two months. Its first product, shown in February, hard-wires Meta's Llama 3.1 8B.

AMD's AI group head Vamsi Boppana framed the purchase as filling out a full-stack platform, with Taalas technology entering the accelerator roadmap alongside Instinct GPUs. Taalas co-founder and chief executive Ljubisa Bajic — previously a founder of Tenstorrent — said the acquisition supplies scale and engineering resources.

Why it matters

Inference, not training, is now where the industry's compute spending is heading, and it is a workload with a very different shape: the same weights served billions of times. Hard-wiring those weights trades all flexibility for efficiency, and until now that trade has looked too brittle to take, because frontier models are replaced every few months. Taalas's two-month turnaround is the bet that the trade has become viable. For AMD, it is a hedge that does not depend on beating Nvidia at general-purpose GPUs: if a meaningful share of production inference settles onto a handful of stable open-weight models, fixed-function silicon could undercut GPU serving economics by a wide margin. It also gives the open-weights ecosystem a hardware argument — a model has to be widely deployed and stable to be worth etching, which favours exactly the models anyone can download.

Sources