NVIDIA Vera: 88 Olympic Arm Cores for AI Agents
NVIDIA
NVIDIA unveils Vera CPU with 88 custom Arm-based 'Olympus' cores designed for AI agent workloads. The chip, part of the Vera Rubin platform, focuses on high IPC, branch prediction, and spatial multithreading to excel in irregular, latency-sensitive tasks.
NVIDIA has announced the Vera CPU, built on the Vera Rubin platform, featuring 88 custom Arm-based 'Olympus' cores using the ARMv9.2 ISA. The Olympus core is engineered for AI agent workloads, emphasizing high instructions-per-clock (IPC) for both single-thread and concurrent tasks. It includes advanced branch prediction, a 10-instruction-wide decoder, and deep out-of-order execution to handle irregular code with many branches and pointer-heavy data structures. NVIDIA uses spatial multithreading (SMT) instead of traditional simultaneous multithreading to divide core resources between two hardware threads, allowing flexible performance and density. The core features eight simple ALUs, six 128-bit SVE2 vector blocks with FP8 support, and a cache hierarchy with 64KB L1i, 96KB L1d, and 2MB L2 per core, along with a 164MB shared L3 cache. The chip integrates a Scalable Coherency Fabric (SCF) with 3.4TB/s bisectional bandwidth, connecting cores, L3, memory controllers, and I/O. Vera supports up to 88 PCIe 6.4 lanes with CXL 3.1, and features Arm CCA, RME-DA, and RME-CDA for security, plus enhanced RAS. In performance comparisons, NVIDIA claims the Olympus core provides 1.9x higher IPC than AMD's Zen 5, with up to 2.3x more branch predictions per cycle and 2.4x more instruction fetches per cycle. AMD has stated that its upcoming EPYC Venice processors will be faster than Vera.
- Abbreviations
- ISA = Instruction Set Architecture — Архитектура набора команд
- IPC = Instructions Per Cycle — Инструкций за такт
- SMT = Simultaneous Multithreading — Одновременная многопоточность
- SVE2 = Scalable Vector Extension 2 — Масштабируемое векторное расширение 2
- FP8 = Floating Point 8-bit — Числа с плавающей запятой 8-битные
- ALU = Arithmetic Logic Unit — Арифметико-логическое устройство
- L1i = Level 1 Instruction Cache — Кэш инструкций 1 уровня
- L1d = Level 1 Data Cache — Кэш данных 1 уровня
- L2 = Level 2 Cache — Кэш 2 уровня
- L3 = Level 3 Cache — Кэш 3 уровня
- SCF = Scalable Coherency Fabric — Масштабируемая когерентная фабрика
- PCIe = Peripheral Component Interconnect Express — Периферийный компонентный интерконнект Express
- CXL = Compute Express Link — Compute Express Link
- CCA = Confidential Compute Architecture — Архитектура конфиденциальных вычислений
- RME-DA = Realm Management Extension - Device Assignment — Расширение управления Realm - назначение устройств
- RME-CDA = Realm Management Extension - Component Device Assignment — Расширение управления Realm - назначение компонентных устройств
- RAS = Reliability, Availability, Serviceability — Надёжность, доступность, обслуживаемость
- MPKI = Misses Per Kilo Instructions — Промахов на тысячу инструкций
- NUMA = Non-Uniform Memory Access — Неоднородный доступ к памяти
- MPAM = Memory System Resource Partitioning and Monitoring — Разделение и мониторинг ресурсов памяти
- NVLink-C2C = NVLink Chip-to-Chip — NVLink между чипами
- SOCAMM2 = Substrate On Chip Advanced Memory Module 2 — Подложка на кристалле усовершенствованный модуль памяти 2
- PSP = Peak System Performance — Пиковая системная производительность
- CBRN = Chemical, Biological, Radiological, Nuclear — Химическая, биологическая, радиологическая, ядерная
Source: ServerNews —
original
