AMD and Cerebras team up for fast AI inference
Cerebras and AMD announced a new partnership that pairs Cerebras's Wafer-Scale Engine 2 chips with AMD's MI300X GPUs. The goal is straightforward: serve AI models with less latency and higher throughput than either could do alone.
The WSE-2 is a 46,000 mm² chip with 940,000 cores and 4 GB of on-chip SRAM per core — no DRAM bottleneck. AMD's MI300X carries 192 GB of HBM3e. Together they cover different parts of the inference workload: the WSE-2 handles the compute-heavy passes, the MI300X manages the memory-heavy ones. The partnership ships as a combined inference platform.
Why this matters for us: la gente que corre modelos locales en casa se beneficia cuando chips como el WSE-2 y el MI300X compiten por precio y rendimiento — más opciones, menos dependencia de un solo proveedor.
“The WSE-2 is a 46,000 mm² chip with 940,000 cores — no DRAM bottleneck.”