AMD acquired Canadian AI chip startup Taalas to deepen AI inference capabilities and challenge NVIDIA dominance.
Advanced Micro Devices (AMD) announced on Thursday that it will acquire Taalas, an innovative AI chip start-up based in Toronto, in a strategic move to challenge NVIDIA's dominance in the AI hardware market. Financial details of the acquisition have not been disclosed.
Founded in 2023, Taalas has developed a revolutionary approach to AI inference that differs fundamentally from traditional GPUs and dataflow architectures. The company's technology directly etches model weights into silicon rather than storing them in high-bandwidth memory (HBM). These model-specific integrated circuits (MSICs) significantly improve inference performance by eliminating the latency and energy costs associated with external memory access.
The company's breakthrough came into public view in February when it unveiled its first test chip, the HC1, built on TSMC's 6nm process technology. The HC1 demonstrated remarkable capabilities, serving Meta's Llama 3.1 8B model at 16,960 tokens per second—48 times faster than NVIDIA's GPUs and 8.5 times faster than Cerebras's accelerators at the time of its unveiling.
For its second-generation HC2 chip, Taalas plans to increase parameter capacity to 20 billion through pipeline parallelism across multiple accelerators. AMD's rack-scale compute platform and in-house system design team position the company to easily scale this architecture. The integration strategy likely involves pairing Instinct-based Helios racks with Taalas-based accelerators, with compute-intensive prompt processing handled by GPUs while token generation is offloaded to the specialized silicon.
The acquisition carries significant implications for AI model development. Developers have been addressing hallucinations through test-time scaling, a technique that trades inference time for accuracy by consuming additional tokens. This approach increases costs and delays chatbot and code assistant responses. If AMD's acquisition enables substantial reductions in token costs while increasing output speeds, it could incentivize developers to extend reasoning times even further, reshaping how AI systems balance accuracy and performance.