Technology·

AMD, Cerebras Enter Partnership to Combine Helios and Wafer Scale Engine

AMD, Cerebras Enter Partnership to Combine Helios and Wafer Scale Engine

By separating inference processing across the two compute platforms, the companies say the architecture can deliver up to 5x higher tokens per second per watt than a Cerebras Wafer-Scale Engine-only configuration.

AMD and Cerebras also announced a partnership to deliver a disaggregated AI inference system combining AMD's Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine. By separating inference processing across the two compute platforms, the companies say the architecture can deliver up to five times higher tokens per second per watt than a Cerebras Wafer-Scale Engine-only configuration. The claim is based on July 2026 modelling by AMD Performance Labs and Cerebras using the Kimi 2.6 1T model at a comparable interactivity point. Lisa Su, AMD CEO, said, “Together with Cerebras, we are extending that leadership into the most latency-sensitive applications and creating a powerful new platform for real-time agentic AI.” “The demand for ultra-fast inference is growing at an unprecedented pace. Cerebras delivers the world’s fastest, ultra-low-latency inference,” said Cerebras CEO Andrew Feldman. Cerebras plans to deploy AMD Helios systems in its data centres, with the joint solution expected to become available initially through Cerebras Cloud in the second half of this year. Cerebras is an AI chipmaker best known for its Wafer-Scale Engine (WSE), a processor built from an entire silicon wafer rather than individual chips. It develops AI hardware and systems for training and inference, positioning its wafer-scale architecture as a faster and more efficient alternative to large GPU clusters. The company went public on the NASDAQ on May 14, after raising more than $2 billion in private funding over roughly 10 years since its founding in 2015. Its final private funding round was a $1 billion Series H completed in February, valuing the company at approximately $23 billion. Cerebras has also secured partnerships with two of the biggest names in AI infrastructure. In January, OpenAI signed a multi-year agreement to deploy 750 MW of Cerebras wafer-scale AI systems to power low-latency inference for ChatGPT and other services, with the rollout planned through 2028. Two months later, Amazon Web Services (AWS) announced that it would deploy Cerebras systems in its data centres and make them available through Amazon Bedrock. The collaboration combines AWS Trainium chips for the prefill stage of inference with Cerebras’ Wafer-Scale Engine for decoding, with the companies claiming the architecture can deliver up to 5× higher token throughput in the same hardware footprint.

This is a summary. Read the full article at the original source.

Read full article at analyticsindiamag