AMD and Cerebras Partner for Low-Latency AI Inference Solutions
AMD and Cerebras are collaborating to develop new AI inference solutions, targeting ultra-low latency applications. The partnership aims to significantly enhance performance and efficiency in demanding AI workloads.

Semiconductor manufacturer AMD and artificial intelligence company Cerebras have announced a joint development effort for AI inference solutions. This collaboration specifically targets use cases requiring ultra-low latency.
The new solution will integrate AMD's Helios rack-level offering with Cerebras' Wafer-Scale Engine. This move is positioned to compete with existing integrated solutions in the market, such as those offered by Nvidia and Groq.
According to the companies' stated roadmaps, AMD's Helios system will function as a high-performance, scalable throughput engine, handling prompt processing and large context windows. Cerebras' Wafer-Scale Engine is designed for the memory-bandwidth-intensive token generation and decoding phases, emphasizing its low-latency capabilities.
When using two compute engines, the combined solution is projected to deliver a fivefold increase in tokens processed per second per watt. Cerebras highlights its Wafer-Scale Engine's architecture, which features hundreds of thousands of compute cores and tens of gigabytes of SRAM on a single wafer, aiming to reduce communication bottlenecks.