AMD and Cerebras Join Forces to Develop Ultra-Low Latency AI Inference Platform
Cerebras Systems
NVIDIA
AMD has partnered with Cerebras Systems to build a computing platform that combines AMD Instinct accelerators with Cerebras' SRAM-based chips to achieve ultra-low latency inference for AI agents. The joint solution aims to fill a gap in AMD's portfolio after Nvidia acquired Groq for $20 billion in December.
AMD and Cerebras Systems are collaborating on a new AI inference platform that merges AMD Instinct accelerators with Cerebras' Wafer Scale Engine (WSE) chips, which use on-chip SRAM memory instead of HBM4, enabling much faster memory access. The partnership addresses a gap in AMD's portfolio after Nvidia's $20 billion acquisition of Groq in December. Cerebras CEO Andrew Feldman criticized Nvidia, comparing it to arms dealers. The combined system handles complex queries on AMD Instinct accelerators and delegates token generation to Cerebras WSE, achieving five times more tokens per watt. While specific metrics are not yet disclosed, the solution is expected to be available in Cerebras Cloud later this year. The platform competes with Nvidia's Vera Rubin paired with Groq 3's LPUs, but AMD and Cerebras claim their system requires only a few dozen chips for a 1-trillion-parameter model like Kimi K2.5, whereas Groq would need two thousand chips.
- Сокращения
- SRAM = Static Random-Access Memory — статическая память с произвольным доступом
- HBM4 = High Bandwidth Memory 4 — высокопропускная память четвёртого поколения
- WSE = Wafer Scale Engine — движок масштаба пластины
- LPU = Language Processing Unit — языковой процессор
Source: 3DNews —
original
