Artificial Intelligence

NVIDIA Unveils Breakthrough AI Inference Economics and Preview Results for Vera Rubin NVL72 in MLPerf Inference v6.1

The economics of artificial intelligence inference are undergoing a fundamental transformation, driven by an intricate interplay of system performance, efficient infrastructure scaling, and continuous software optimization. As enterprises and hyperscalers race to deploy generative AI applications at a global scale, the underlying hardware and software ecosystems must evolve in tandem to keep operational costs manageable while maximizing revenue-generating potential. High system performance directly translates to a greater volume of generated tokens, which inherently boosts top-line revenue. Meanwhile, efficient infrastructure scaling guarantees that throughput increases proportionally as additional hardware resources are integrated into a data center, drastically reducing the cost per served user. Finally, continuous software optimization squeezes maximum utility out of capital-intensive infrastructure investments, ensuring that computational capacity does not go underutilized.

At the heart of this operational trifecta lies platform fungibility—the architectural capability of a single, unified infrastructure to seamlessly execute any AI model or workload. From intensive training sessions to real-time inference, from complex recommender systems to multi-step reasoning pipelines, and from natural language processing to high-definition video generation, fungibility ensures that data center utilization remains remarkably high. NVIDIA’s accelerated computing architecture is purpose-built to address these complex requirements, a fact vividly underscored by the newly released MLPerf Inference v6.1 benchmark results. These benchmarks offer critical insights for organizations tasked with designing long-term AI infrastructure strategies where performance, scalability, and software velocity dictate ultimate market success.

Preview Performance and the Rise of the Vera Rubin Architecture

Marking a major milestone in high-performance computing, NVIDIA submitted preview results for its next-generation Vera Rubin NVL72 platform during the MLPerf Inference v6.1 evaluation cycle. The submissions focused heavily on two of the industry’s most computationally demanding benchmarks: the DeepSeek-R1 reasoning model and the Qwen3-VL vision-language model. These workloads represent the cutting edge of AI complexity, demanding exceptional memory bandwidth, sophisticated orchestration, and massive parallel processing capabilities.

The Vera Rubin NVL72 architecture delivered staggering performance gains during testing. Across offline, server, and interactive scenarios on the Qwen3-VL benchmark, the Vera Rubin platform achieved up to 3.7 times higher throughput than the preceding GB300 NVL72 system. This remarkable leap was achieved by leveraging vLLM in conjunction with the open-source NVIDIA Dynamo inference framework. Similarly, on the DeepSeek-R1 benchmark, utilizing the NVIDIA TensorRT-LLM library, Vera Rubin delivered up to 2.5 times higher throughput than its predecessor. These early preview metrics not only highlight NVIDIA’s relentless pace of innovation but also signal how much further performance is expected to climb as continuous software refinements mature ahead of full commercial availability.

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

From a financial and operational perspective, these performance figures carry profound implications. Each Vera Rubin NVL72 rack is engineered to output significantly more tokens, service a vastly expanded user base, and generate substantially higher revenue streams than a comparable GB300 NVL72 rack, all while driving down the overall cost per token.

Full-Stack Co-Design and Advanced Engineering

The extraordinary performance metrics recorded in the MLPerf Inference v6.1 suite are not the result of hardware improvements alone; rather, they are a testament to rigorous full-stack co-design spanning both silicon and software engineering. The Vera Rubin architecture incorporates enhanced Tensor Cores and a redesigned Transformer Engine specifically optimized to accelerate both the prefill and decode stages of the inference lifecycle. Furthermore, the integration of NVFP4 numerical precision significantly reduces the memory footprint across model weights, attention mechanisms, and Key-Value (KV) caches, thereby boosting overall throughput without triggering a noticeable degradation in output quality.

To handle the immense computational load of Mixture-of-Experts (MoE) architectures—which underpin complex models like DeepSeek-R1 and Qwen3-VL—Vera Rubin’s engineering team heavily utilized disaggregated serving techniques. By separating the prefill and decode phases and pairing them with large-scale expert parallelism, the platform achieves unprecedented efficiency.

This hardware-software synergy is anchored by the NVL72 scale-up domain, which relies on sixth-generation NVIDIA NVLink and dedicated NVLink Switches. This networking fabric delivers a phenomenal tenfold increase in packet rates and three times lower latency compared to off-the-shelf Ethernet alternatives. This robust interconnect foundation is what enables advanced parallelism and disaggregation techniques to operate smoothly at a massive rack scale. Proof of this ecosystem readiness was further demonstrated by industry partners such as Nebius, which also submitted Vera Rubin NVL72 preview results and showcased stellar performance metrics.

The Evolution of Inference: Transitioning to Agentic AI Workloads

As artificial intelligence rapidly transitions from simple text-generation chatbots to sophisticated autonomous agents capable of reasoning, planning, and executing multi-step workflows, the metrics used to measure inference performance must also evolve. Traditional throughput benchmarks, while still essential, fail to capture the multi-faceted nature of agentic AI.

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

Addressing this paradigm shift, preview testing on benchmarks such as the SemiAnalysis AgentX framework revealed that the Vera Rubin NVL72 platform delivers up to 30 times better performance than the GB300 NVL72. Recognizing the industry-wide shift toward autonomous workflows, upcoming evaluation suites like the MLPerf Endpoints benchmark are expected to introduce standardized measurements specifically tailored for agentic inference workloads. This ensures that infrastructure evaluations remain aligned with the real-world demands of modern enterprise applications.

Scalability and Rack-Level Efficiency of the GB300 NVL72

While preview architectures capture future potential, current deployments rely heavily on proven scalability. Scaling efficiency—defined as how effectively the addition of more graphical processing units translates into proportional throughput gains—remains a primary indicator of infrastructure productivity. NVIDIA achieves high scaling efficiency through a combination of high-bandwidth, low-latency scale-up interconnects within individual racks, high-speed networking between separate racks, and intelligent request orchestration across distributed nodes.

During the MLPerf Inference v6.1 evaluations, NVIDIA’s DeepSeek-R1 submission demonstrated remarkable elasticity, scaling seamlessly from a single GB300 NVL72 rack containing 72 GPUs up to a massive four-rack configuration comprising 288 GPUs. In the offline scenario, the system achieved an astonishing 99% scaling efficiency, meaning that throughput grew almost in direct proportion to the massive influx of added hardware.

Achieving high scaling efficiency is notoriously difficult in distributed computing. Without careful architectural alignment, adding double the hardware often results in diminishing returns due to communication bottlenecks, making infrastructure costs vastly outpace performance returns. The flawless scaling of the GB300 NVL72 proves that hardware, interconnect fabrics, and orchestration software can operate as a unified, cohesive entity. Furthermore, the GB300 NVL72 exhibited exceptional rack-scale efficiency on the WAN 2.2 text-to-video generation benchmark. It achieved 0.65 frames of 720p video per second with an average generation time of 5.7 seconds per video—delivering nine times higher throughput and 7.5 times lower latency than a single isolated compute node.

Software Velocity and Continuous Post-Submission Gains

NVIDIA’s development methodology relies heavily on an aggressive cadence of continuous software enhancements, ensuring that hardware assets appreciate in capability long after initial deployment. In the v6.1 benchmark cycle alone, GB300 NVL72 performance on the Qwen3-VL workload improved by up to 1.6 times compared to the previous v6.0 results. These performance leaps were unlocked through optimizations such as lower KV cache precision, expanded kernel fusion, refined execution kernels, and advanced disaggregated serving via vLLM and NVIDIA Dynamo.

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

Significantly, software engineering efforts did not pause when the MLPerf submission window closed. Post-submission optimizations—though not yet independently verified by MLCommons at the time of publication—have yielded even higher performance metrics on complex workloads like the GPT-OSS-120B language model and DLRMv3 recommendation systems. This continuous software velocity ensures that enterprise customers running NVIDIA infrastructure benefit from compounding performance dividends over the lifecycle of their hardware deployments.

Democratizing AI Infrastructure Across the Ecosystem

The impact of these technological advancements extends far beyond ultra-scale data centers. Demonstrating versatility across the entire computing spectrum, NVIDIA also submitted edge computing results for the Jetson AGX Thor platform. Utilizing NVIDIA TensorRT Edge-LLM on the newly established Edge-Agentic benchmark with the Qwen3.6-27B model, the compact edge processor showcased robust local intelligence capabilities.

Furthermore, the broad participation of NVIDIA’s partner ecosystem in the MLPerf Inference v6.1 round underscores the mainstream enterprise readiness of these platforms. A total of 19 major technology vendors and cloud providers participated, with eight specifically showcasing multi-node Blackwell NVL72 systems. The participating ecosystem includes industry leaders such as ASUS, Microsoft Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, Hewlett Packard Enterprise (HPE), Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure (OCI), Quanta Cloud Technology (QCT), Red Hat, ScitiX, Supermicro, and Wiwynn.

Broad Implications for the Global Technology Landscape

The release of the MLPerf Inference v6.1 results, highlighted by the preview dominance of the Vera Rubin architecture and the proven hyper-scalability of the GB300 NVL72, paints a clear picture of the future trajectory of artificial intelligence. As enterprises transition from proof-of-concept AI experiments to fully autonomous, agentic enterprise deployments, the economic pressures surrounding compute efficiency will only intensify.

By aggressively co-designing full-stack solutions—combining specialized silicon innovations like NVFP4 precision and advanced Tensor Cores with high-speed interconnects like sixth-generation NVLink and continuously optimized software frameworks—NVIDIA is setting a rigorous benchmark for the industry. For CIOs and data center architects navigating the capital-intensive world of generative AI infrastructure, these developments offer a vital roadmap: sustainable, revenue-positive AI economics require deep integration across the entire technology stack, capable of scaling effortlessly from the compact edge to the largest AI factories in the world.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button
Lock It Soft
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.