NVIDIA spent years teaching the market that the GPU was the center of modern computing. Now it is making a second argument: the GPU is not enough.

The company says its first custom CPU, Vera, is shipping at scale. AWS has received an initial Vera server and Vera Rubin GPU system, joining earlier deliveries to Oracle Cloud Infrastructure, Anthropic, OpenAI, and SpaceXAI. Delivery is evidence of hardware entering partner environments, not proof of broad production performance. The distinction matters because NVIDIA is also the source of the performance claims.

The architectural thesis is solid. An agent does more than generate tokens. It retrieves long context, calls tools, runs code in sandboxes, coordinates state, moves data, and waits on external systems. GPUs handle accelerated model work. Much of the surrounding mess lands on CPUs. When thousands of agents operate concurrently, that supporting work can become the bottleneck nobody put on the benchmark slide.

NVIDIA says Vera has 88 custom Olympus cores, 1.2 terabytes per second of memory bandwidth, and up to 1.8 times faster per-core performance on agentic workloads. It also says Vera and Rubin systems can feed GPUs at twice the energy efficiency of traditional infrastructure. Those numbers need independent workload testing, especially across software stacks that do not look like NVIDIA's preferred reference architecture.

The strategic move is larger than one chip. Vera sits beside Rubin GPUs, BlueField data processing units, Spectrum networking, NVLink, and rack designs. NVIDIA is turning what used to be a component sale into control of the whole factory. Better integration can improve utilization. It can also make the customer increasingly dependent on one vendor's roadmap.

Watch the deployment details: which workloads actually migrate, how much CPU time agents consume, whether cloud customers can measure total cost, and how easily the software layer moves across competing architectures. Shipping hardware starts the argument. Production economics decide it.

If agents become a durable computing category, the winner will not merely build the fastest accelerator. It will keep every expensive part of the system busy while the agent is searching, deciding, calling, failing, and trying again. NVIDIA understands the machinery. The market still has to verify the leverage.

LaunchPad positionThe agent economy will not run on GPUs alone. Whoever controls the full path between model, memory, tools, and compute can capture more of the system and remove more of the bottleneck.
Reporting standard

This report draws on the linked primary sources and reputable reporting. Company statements are treated as claims until independently demonstrated.