AMD and Cerebras, usually rivals, team up on a disaggregated AI inference system

AMD and Cerebras Systems, two companies more accustomed to competing for AI accelerator customers than collaborating, announced a technical partnership July 23 that splits an AI inference workload between their respective hardware. The disaggregated design combines AMD’s Helios rackscale system, built around its Instinct GPUs, with Cerebras’s Wafer-Scale Engine, each handling the stage of inference it does best.

The architecture separates inference into two phases. AMD Helios handles prompt processing and large-context compute: the high-throughput stage where the model digests the user’s input and assembles the relevant context. Cerebras’s Wafer-Scale Engine then takes over for token generation, the decode phase, where low latency and high memory bandwidth matter most. Both engines operate as a single inference workflow through a unified API.

AMD and Cerebras claim the combined system delivers up to five times higher tokens per second per watt compared to a Cerebras-only configuration, based on internal modeling with the Kimi 2.6 1-trillion-parameter model. The metric matters: as AI inference scales from chatbots to real-time agents, robotics, and copilots, energy efficiency per token directly determines deployment cost and feasibility.

“AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach,” said AMD CEO Lisa Su in a statement. Cerebras CEO Andrew Feldman framed the partnership as an access play: “Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers.”

If our reporting has earned your trust, consider helping us continue our work.

Back evidence-based news

The partnership is notable because both companies sell hardware designed for the same general use case. AMD’s Instinct GPUs compete directly with Nvidia’s data center lineup, while Cerebras has positioned its wafer-scale chips as an alternative to both, offering extreme single-die performance by building a chip the size of a silicon wafer rather than cutting it into individual dies. The disaggregated approach acknowledges that no single architecture is optimal for both the memory-bandwidth-intensive prompt-processing stage and the latency-sensitive decode stage of modern inference.

Cerebras plans to deploy AMD Helios systems in its own data centers, with the joint solution available first through Cerebras Cloud in the second half of 2026. The announcement was part of AMD’s Advancing AI 2026 event, which also saw the launch of the Helios rackscale system and the Instinct MI400 series of GPUs.

Sources: AMD and Cerebras announce ultra-low-latency AI inference solution (AMD Newsroom, Jul 23, 2026); AMD partners with Cerebras for AI inference system (DCD, Jul 23, 2026); Cerebras press release (Cerebras, Jul 23, 2026)

Scroll to Top