# Nvidia Begins Volume Shipments of Vera, an 88-Core Server CPU Aimed at Agent Workloads

Oracle Cloud Infrastructure plans to deploy hundreds of thousands of the processors starting this year, and early units went to Anthropic, OpenAI and Amazon Web Services.

- Published: 2026-08-28T06:11:21.137Z
- Canonical: https://polylog.news/ai/2026-08-28/nvidia-begins-volume-shipments-of-vera-an-88-core-server-cpu
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [NVIDIA Blog](https://blogs.nvidia.com/blog/vera-cpu-delivery/), [NVIDIA Newsroom](https://nvidianews.nvidia.com/news/nvidia-unveils-vera-the-cpu-for-agents), [VideoCardz](https://videocardz.com/newz/nvidia-details-vera-cpu-with-88-olympus-cores-176-threads-and-1-2-tb-s-lpddr5x-memory)

Nvidia said its Vera central processing unit (CPU) has entered full production and is [shipping at scale](https://blogs.nvidia.com/blog/vera-cpu-delivery/), with Ian Buck, the company's vice president of hyperscale and high-performance computing, hand-delivering the first systems to Amazon Web Services, Oracle Cloud Infrastructure and three model developers, Anthropic, OpenAI and SpaceXAI. Commercial availability through system builders and cloud partners is set for the autumn.

Vera is the first server processor where Nvidia designed the core itself rather than licensing an Arm design. It carries [88 custom Olympus cores and 176 hardware threads](https://videocardz.com/newz/nvidia-details-vera-cpu-with-88-olympus-cores-176-threads-and-1-2-tb-s-lpddr5x-memory) on a monolithic compute die with 164 megabytes of unified last-level cache, up to 1.2 terabytes per second of LPDDR5X memory bandwidth using SOCAMM2 modules, up to 1.5 terabytes of memory per socket, and up to 3.4 terabytes per second of core-to-core bandwidth. Nvidia claims [1.8 times faster task completion than x86 processors](https://nvidianews.nvidia.com/news/nvidia-unveils-vera-the-cpu-for-agents) on agentic AI, reinforcement learning and data-processing workloads, a vendor figure that has not been independently benchmarked across those three categories.

The design choice worth noting is the emphasis on single-thread performance and memory capacity rather than raw core count. Agent workloads spend a large share of wall-clock time on orchestration, tool calls, sandboxed code execution, retrieval and data preparation, all of which run on the host CPU while the accelerator waits. Oracle's stated plan to field hundreds of thousands of Vera parts beginning this year is a bet that host-side serialization, not GPU throughput, is what limits agent throughput at scale.

## What this means

Nvidia is extending its reach from accelerators into the host processor socket, which puts direct pressure on Intel and AMD in the one part of the AI server rack they still reliably supply. The mechanism is bundling. If the Vera-Rubin platform makes the CPU and GPU a single procurement decision, x86 vendors lose the AI data center as a growth market and are left defending general-purpose enterprise servers. What decides this is whether the 1.8-times agentic advantage survives third-party testing on real orchestration workloads rather than the workloads Nvidia chose to test.

## What to watch

- Independent benchmark suites running agent orchestration and reinforcement-learning data pipelines on Vera against current x86 parts, which would confirm or deflate the headline speedup.
- Whether AMD and Intel respond with server parts tuned for host-side agent work rather than core count, which would show they accept Nvidia's framing of the workload.
- Oracle's actual deployed Vera volume against its stated plan, which is the clearest read on whether cloud buyers will pay for a non-x86 host.
