The speed result is dramatic. Taalas' HC1 test chip, fabricated on TSMC's 6nm process, served Meta's Llama 3.1 8B at 16,960 tokens per second in February 2026 — described as up to 48x faster than an equivalent Nvidia GPU inference. Some sources cite up to 73x faster depending on the specific Nvidia baseline.
Because the weights are physically baked into the silicon at the factory, each Taalas chip can run only the single model it was manufactured for. You cannot load a different model onto it.
However, Taalas developed a toolchain that mitigates this rigidity. It finalizes only the top two metal layers on a roughly 100-layer chip stack, enabling a roughly two-month turnaround to tape out a new model. The underlying base wafer design stays constant, making re-spins fast compared to traditional ASIC development.
AMD plans to integrate Taalas MSICs into a disaggregated, heterogeneous compute architecture alongside its Instinct GPUs and Helios rack-scale systems.
In this design, general-purpose Instinct GPUs handle training and flexible inference workloads, while Taalas chips are deployed as high-throughput, low-latency dedicated inference accelerators for the most heavily demanded models. The Helios platform connects them in a unified rack-scale fabric.
This approach lets AMD offer customers a tiered inference solution: flexible GPUs for diverse workloads and ultra-efficient MSICs for fixed, high-volume models.
Deal terms were undisclosed. The acquisition is part of a string of major AMD purchases: the $4.9 billion acquisition of ZT Systems (July 2024) and the $665 million acquisition of Silo AI (August 2024).
On the competitive front, the deal positions AMD against Nvidia's reported $20 billion licensing deal with Groq (announced earlier in 2026), which gives Nvidia access to Groq's LPU inference architecture. AMD's bet is that Taalas' model-specific hardwiring can deliver even better performance per watt for fixed, high-volume inference workloads.
Taalas was founded in 2023 in Toronto by Ljubisa Bajic (formerly CEO of Tenstorrent, another AI chip startup) and a team that included ex-AMD engineers.
The company raised $219 million in venture funding from investors including Eclipse Ventures, Makers Fund, Radical Ventures, and others before being acquired.
It demonstrated the HC1 chip in February 2026 running Llama 3.1 8B at the 16,960 tokens/second benchmark that caught AMD's attention.
AMD's Taalas acquisition represents a bet that the future of AI inference is not in ever-faster general-purpose GPUs but in purpose-built chips that sacrifice flexibility for extreme efficiency on the models that matter most. By pairing Taalas MSICs with Instinct GPUs in a disaggregated architecture, AMD is building a system that can handle both the long tail of diverse models and the high-volume traffic of the most popular ones — a direct challenge to Nvidia's GPU-dominated inference stack.