NVIDIA says its new Vera Rubin NVL72 AI system can deliver up to 30x higher throughput per megawatt than the company’s GB300 NVL72 on agentic AI workloads. The early results come from tests using real-world coding sessions and point to lower energy use and token costs for AI systems that perform long, multi-step tasks.
The results were measured using the SemiAnalysis AgentX workload, which uses recorded coding sessions with real tool calls, growing context and sub-agents. NVIDIA says Vera Rubin NVL72 also delivers up to 35x lower cost per million tokens than GB300 NVL72 in the tested workload.
Agentic AI uses much more computing than a simple chatbot request. An AI agent may search databases, call tools, ask other AI agents for help and repeatedly reason through a problem before producing a final answer.
READ ALSO: https://modernmechanics24.com/post/el-nino-oceans-record-21-1c-early-peak/
This creates a major challenge for AI data centers. Each step can add more information to the next step, causing the amount of context to grow significantly. NVIDIA’s new system is designed to process these long and complex workloads more efficiently.
Vera Rubin NVL72 uses several hardware and software techniques to improve inference performance. These include separate processing for context and response generation, large-scale expert parallelism, distributed KV caching and routing that sends requests to GPUs where useful cached information is already available.
The system also uses NVIDIA’s fifth-generation Tensor Cores, third-generation Transformer Engine and NVFP4 quantization. NVIDIA says its latest NVLink technology provides faster communication between GPUs, helping the system handle large AI models and long-context workloads.
According to NVIDIA, its DSX MaxLPS technology can manage power across GPUs, racks and workloads. The company says this could allow up to 40% more GPUs to operate within the same power budget at AI factory scale.
The results have some limits. The Vera Rubin figures are early measurements from NVIDIA and are currently pending SemiAnalysis review. The tests also do not yet include Vera CPU performance for tool-calling workloads, and future software improvements could change performance on both Vera Rubin and GB300.
WATCH ALSO: https://modernmechanics24.com/post/lockheed-martin-grizzlymissile-launcher/
The efficiency gains could become important as companies deploy AI agents for coding, research, customer service and other complex tasks. If independent testing confirms NVIDIA’s results, Vera Rubin NVL72 could help data centers run more agentic AI work within existing power limits while lowering the cost of generating AI tokens.














One Response