AWS and Cerebras Announce Partnership for Ultra-Fast AI Inference on Amazon Bedrock
- AWS and Cerebras will deploy CS-3 systems in AWS data centers for the fastest AI inference on Amazon Bedrock, launching in coming months[1][4].
- Disaggregated architecture uses AWS Trainium for prefill and Cerebras CS-3 for decode, connected by Elastic Fabric Adapter, offering order-of-magnitude speed gains[1][2].
- Supports open-source LLMs and Amazon Nova models for real-time applications like coding assistance[1][4].
- Cerebras CS-3 provides up to 25 times faster inference than Nvidia GPUs in decode stage[3].
- Cerebras valued at $23 billion after $1 billion funding in February 2026[3][5]
Amazon Web Services and Cerebras Systems announced a collaboration to deploy Cerebras CS-3 systems in AWS data centers, delivering high-speed AI inference for generative AI workloads through Amazon Bedrock. The partnership introduces a disaggregated architecture combining AWS Trainium for prefill and Cerebras CS-3 for decode, connected via Elastic Fabric Adapter networking, with availability expected in the coming months[1][4].
Partnership Details
The collaboration makes AWS the first cloud provider to offer Cerebras's disaggregated inference solution exclusively through Amazon Bedrock. AWS Trainium servers handle the prefill stage, while Cerebras CS-3 systems manage decode, linked by Elastic Fabric Adapter. This setup promises up to 5x higher token capacity and inference speeds an order of magnitude faster than current options[1][2][4]. David Brown, Vice President of Compute & ML Services at AWS, stated, “Inference is where AI delivers real value to customers, but speed remains a critical bottleneck for demanding workloads like real-time coding assistance and interactive applications”[1].
Performance and Applications
Cerebras CS-3 is described as the world's fastest AI inference system, with thousands of times greater memory bandwidth than the fastest GPUs. It excels in reasoning models that generate more tokens, supporting applications such as agentic coding used by OpenAI, Cognition, and Mistral. Internal testing shows inference under 10 milliseconds for billion-parameter models, 10 to 100 times faster than GPU clusters[1][2]. Later this year, AWS will support leading open-source LLMs and Amazon Nova models on this hardware[1].
Cerebras Background
The multi-year partnership follows Cerebras's momentum, including a January 2026 deal with OpenAI valued over $10 billion for up to 750 megawatts of compute and a $1 billion funding round in February 2026 that raised its valuation to about $23 billion[3][5][6]. Cerebras claims its Wafer-Scale Engine delivers up to 25 times faster performance than Nvidia GPUs in the decode stage of inference[3]. CEO Andrew Feldman noted, “Partnering with AWS to build a disaggregated inference solution will bring the fastest inference to a global customer base”[1].