AMD has acquired Taalas, an AI chip startup based in Toronto, to challenge the current landscape of inference hardware. Taalas uses a unique approach where model weights are etched directly into the silicon of the chip. This method bypasses the standard reliance on high-bandwidth memory for storing model data.
The resulting hardware functions as a model-specific integrated circuit. In initial testing with the HC1 chip, Taalas demonstrated performance levels significantly faster than existing industry standards. By hard-coding the model weights into the processor, the system reduces the power and space requirements compared to traditional GPU clusters. This approach offers a potential advantage for large-scale AI deployment, as AMD aims to integrate this technology into its existing rack-scale compute platforms.
While the performance gains are substantial, the technology introduces a trade-off regarding flexibility. Because the model weights are physically etched onto the chip, updating to a completely new AI model requires a re-spin of the hardware components. However, Taalas suggests that this process is far less expensive and faster than a full manufacturing restart, as only specific layers of metal require changes.
AMD expects this technology to appeal to infrastructure providers and large model developers who prioritize consistent, high-speed inference. By pairing this silicon with its Instinct-based Helios racks, AMD intends to provide a hybrid architecture where compute-heavy tasks occur on GPUs while token generation is offloaded to the Taalas-based accelerators. The deal is expected to close in the fourth quarter of this year.

