NVIDIA Unveils Breakthrough AI Inference Economics and Preview Results for Vera Rubin NVL72 in MLPerf Inference v6.1 Benchmarks

The economics of artificial intelligence have fundamentally shifted from a paradigm of raw computational capacity to an intricate balancing act centered on system performance, efficient infrastructure scaling, and continuous software optimization. As enterprises and hyperscalers race to deploy generative AI models at global scales, these three pillars dictate the fundamental return on investment for data centers. Higher system performance directly translates to a greater volume of generated tokens, which in turn drives higher revenue potential. Simultaneously, efficient infrastructure scaling guarantees that throughput grows proportionally with the addition of new hardware, mitigating the risk of diminishing returns and ensuring that organizations require fewer resources to serve massive user bases. Rounding out the triad, continuous software optimization extracts maximum value from existing capital expenditures, squeezing unprecedented performance out of deployed silicon. Beneath all of these operational levers lies the critical concept of platform fungibility—the architectural capability of a single, unified infrastructure to seamlessly run any model, execute any workload spanning from training to inference, handle tasks from recommendation engines to complex reasoning, and process modalities ranging from text to video. By keeping utilization rates perpetually high, fungibility eliminates the idle waste traditionally associated with specialized hardware silos. The NVIDIA accelerated computing platform is purposefully architected to optimize across every single one of these vectors, a design philosophy vividly highlighted by the MLPerf Inference v6.1 benchmark results released today by MLCommons.
For enterprise decision-makers tasked with shaping multi-million-dollar AI infrastructure roadmaps, system performance, scaling efficiency, and software velocity are no longer secondary metrics; they are the primary determinants of long-term inference economics. Organizations are increasingly wary of hardware that depreciates rapidly or fails to scale linearly. In response to these market demands, NVIDIA submitted preview results for its upcoming Vera Rubin NVL72 architecture on two of the most computationally demanding and structurally complex benchmarks within the MLPerf Inference v6.1 suite: DeepSeek-R1, a leading reasoning model utilizing mixture-of-experts architecture, and Qwen3-VL, a state-of-the-art vision-language model.
The preview performance figures released for the Vera Rubin NVL72 architecture underscore a dramatic acceleration in NVIDIA’s hardware-software co-design cadence. Across offline, server, and interactive operational scenarios on the Qwen3-VL benchmark, the Vera Rubin NVL72 delivered up to 3.7 times higher throughput than the preceding GB300 NVL72 system. This remarkable leap was achieved utilizing vLLM integrated with the NVIDIA Dynamo open-source inference framework. Meanwhile, on the DeepSeek-R1 reasoning benchmark, utilizing the NVIDIA TensorRT-LLM library, the Vera Rubin architecture achieved throughput up to 2.5 times higher than the GB300 NVL72. These early preview results not only signal the formidable capabilities of the upcoming platform but also illustrate how performance will continue to compound over time through continuous software optimizations long before the hardware reaches general commercial availability.
The economic implications of these performance multipliers are profound. Each Vera Rubin NVL72 rack is engineered to deliver a significantly larger volume of tokens, serve a vastly expanded user base, and generate substantially higher revenue streams than a comparable GB300 NVL72 rack, all while driving down the total cost per token. This efficiency is the direct byproduct of comprehensive full-stack co-design spanning both silicon and software layers. The Vera Rubin architecture features enhanced Tensor Cores and an advanced Transformer Engine designed to accelerate both the prefill and decode stages of the inference lifecycle. Furthermore, the adoption of NVFP4 numerical precision significantly reduces the memory footprint across model weights, attention mechanisms, and key-value (KV) caches, enabling higher throughput with negligible degradation in output fidelity.
To sustain this performance, NVIDIA’s Vera Rubin submissions made extensive use of disaggregated serving—an architectural technique that physically or logically separates the prefill and decode stages—alongside large-scale expert parallelism. This combination maximizes operational efficiency across the complex mixture-of-experts (MoE) layers that underpin advanced reasoning and multimodal models like DeepSeek-R1 and Qwen3-VL. The scale-up domain of the NVL72 platform, powered by sixth-generation NVIDIA NVLink and NVLink switches, provides an indispensable interconnect foundation. Delivering tenfold higher packet rates and threefold lower latency compared to off-the-shelf Ethernet fabrics, this high-bandwidth interconnect is what makes sophisticated rack-scale parallelism practically viable. This hardware-software synergy extends outward into NVIDIA’s robust partner ecosystem, exemplified by cloud provider Nebius, which also submitted Vera Rubin NVL72 preview results, demonstrating exceptional performance out of the gate.

The Evolution of Inference Benchmarks in the Age of AI Agents
The metrics by which the industry measures inference performance are undergoing a structural evolution, driven largely by the proliferation of autonomous AI agents capable of reasoning, planning, and executing multi-step workflows. Traditional throughput benchmarks, while valuable for measuring raw token generation, fail to capture the nuanced latency and decision-making bottlenecks inherent in agentic workflows. Recognizing this paradigm shift, industry benchmarks such as the SemiAnalysis AgentX have emerged to measure end-to-end agentic performance. In preview testing on these specialized benchmarks, the Vera Rubin NVL72 delivered up to 30 times better performance than the GB300 NVL72, signaling a monumental leap forward for enterprise automation and autonomous system deployment. Furthermore, the upcoming launch of the MLPerf Endpoints benchmark is expected to bring standardized, rigorous measurement to agentic inference workloads, providing the industry with a unified yardstick for evaluating complex reasoning systems.
NVIDIA GB300 NVL72 Scales with Leading Efficiency
While breakthrough single-node and rack performance captures headlines, scaling efficiency—defined as how effectively the addition of supplementary GPUs translates into proportional throughput gains—remains the ultimate litmus test for data center productivity. NVIDIA addresses this challenge through a combination of high-bandwidth, low-latency scale-up interconnects within individual racks, high-speed networking fabrics between separate racks, and intelligent request orchestration across distributed nodes.
In MLPerf Inference v6.1 evaluations, NVIDIA’s DeepSeek-R1 submission demonstrated exceptional linear scalability, expanding seamlessly from a single GB300 NVL72 rack housing 72 GPUs to a four-rack configuration comprising 288 GPUs. In this multi-rack deployment, the system achieved a staggering 99% scaling efficiency in the offline scenario, meaning that throughput grew almost entirely in direct proportion to the massive influx of additional hardware.
This level of scaling efficiency is critical for modern data center economics. Simply adding more GPUs to a cluster does not automatically guarantee proportional throughput gains; if doubling the hardware footprint yields only single-digit performance improvements, the capital expenditure required to scale out the infrastructure quickly outpaces any financial return. Achieving true linear scaling requires deep, foundational harmony among the silicon architecture, the network interconnect fabric, and the orchestration software layer.

The scaling prowess of the GB300 NVL72 was similarly demonstrated on demanding multimodal workloads, such as the WAN 2.2 text-to-video generation benchmark. In these tests, the platform achieved 0.65 720p videos per second with an average generation time of just 5.7 seconds per video. This represented a ninefold increase in throughput and a 7.5-fold reduction in latency compared to a single-node configuration, proving that rack-scale architecture can effectively handle heavy, compute-intensive generative media workloads without buckling under pressure.
Software Optimizations Drive Continuous Performance Gains
Hardware is only as capable as the software stack that commands it. NVIDIA’s accelerated computing platform undergoes rigorous, continuous software development, ensuring that deployed systems receive ongoing performance enhancements and feature expansions long after initial installation.
In the transition from MLPerf Inference v6.0 to v6.1, performance for the GB300 NVL72 running the Qwen3-VL benchmark jumped by up to 1.6x. These gains were unlocked through a combination of lowered KV cache precision, additional kernel fusion, optimized execution kernels, and refined disaggregated serving implementations utilizing vLLM and NVIDIA Dynamo.
Crucially, software velocity did not halt with the official v6.1 submission deadlines. Post-submission evaluations—which have not yet been formally verified by MLCommons—demonstrate even further performance gains on large language models like GPT-OSS-120B and recommendation systems like DLRMv3, illustrating that enterprise customers continue to extract greater value from their hardware investments over time through routine software updates.
Ecosystem Participation and Deployment at Every Scale

The reach of NVIDIA’s inference platform extends far beyond high-end data center racks, encompassing edge devices and distributed enterprise environments alike. Alongside its data center announcements, NVIDIA submitted results for the Jetson AGX Thor platform utilizing NVIDIA TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark powered by Qwen3.6-27B, showcasing low-power, high-performance edge reasoning capabilities.
This technological momentum is supported by a broad, highly engaged global partner ecosystem. Nineteen major technology vendors—including eight deploying multi-node Blackwell NVL72 systems—participated in the MLPerf Inference v6.1 round, demonstrating immediate market readiness and interoperability. The participating ecosystem includes ASUS, Microsoft Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, Hewlett Packard Enterprise (HPE), Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure (OCI), Quanta Cloud Technology (QCT), Red Hat, ScitiX, Supermicro, and Wiwynn.
From compact edge devices deployed in industrial environments to massive hyperscale AI factories powering global language and video models, NVIDIA continues to push the boundaries of full-stack performance. Driven by an unwavering annual cadence of platform architectures, continuously evolving software libraries, and an expansive global partner network, the industry is well-equipped to deliver scalable, economically viable artificial intelligence to organizations worldwide.







