{"id":7629,"date":"2026-09-18T22:53:26","date_gmt":"2026-09-18T22:53:26","guid":{"rendered":"https:\/\/lockitsoft.com\/?p=7629"},"modified":"2026-09-18T22:53:26","modified_gmt":"2026-09-18T22:53:26","slug":"nvidia-unveils-breakthrough-ai-inference-economics-and-preview-results-for-vera-rubin-nvl72-in-mlperf-inference-v6-1","status":"publish","type":"post","link":"https:\/\/lockitsoft.com\/?p=7629","title":{"rendered":"NVIDIA Unveils Breakthrough AI Inference Economics and Preview Results for Vera Rubin NVL72 in MLPerf Inference v6.1"},"content":{"rendered":"<p>The economics of artificial intelligence inference are undergoing a fundamental transformation, driven by an intricate interplay of system performance, efficient infrastructure scaling, and continuous software optimization. As enterprises and hyperscalers race to deploy generative AI applications at a global scale, the underlying hardware and software ecosystems must evolve in tandem to keep operational costs manageable while maximizing revenue-generating potential. High system performance directly translates to a greater volume of generated tokens, which inherently boosts top-line revenue. Meanwhile, efficient infrastructure scaling guarantees that throughput increases proportionally as additional hardware resources are integrated into a data center, drastically reducing the cost per served user. Finally, continuous software optimization squeezes maximum utility out of capital-intensive infrastructure investments, ensuring that computational capacity does not go underutilized.<\/p>\n<p>At the heart of this operational trifecta lies platform fungibility\u2014the architectural capability of a single, unified infrastructure to seamlessly execute any AI model or workload. From intensive training sessions to real-time inference, from complex recommender systems to multi-step reasoning pipelines, and from natural language processing to high-definition video generation, fungibility ensures that data center utilization remains remarkably high. NVIDIA\u2019s accelerated computing architecture is purpose-built to address these complex requirements, a fact vividly underscored by the newly released MLPerf Inference v6.1 benchmark results. These benchmarks offer critical insights for organizations tasked with designing long-term AI infrastructure strategies where performance, scalability, and software velocity dictate ultimate market success.<\/p>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_82_2 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<span class=\"ez-toc-title-toggle\"><a href=\"#\" class=\"ez-toc-pull-right ez-toc-btn ez-toc-btn-xs ez-toc-btn-default ez-toc-toggle\" aria-label=\"Toggle Table of Content\"><span class=\"ez-toc-js-icon-con\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/span><\/a><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/lockitsoft.com\/?p=7629\/#Preview_Performance_and_the_Rise_of_the_Vera_Rubin_Architecture\" >Preview Performance and the Rise of the Vera Rubin Architecture<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/lockitsoft.com\/?p=7629\/#Full-Stack_Co-Design_and_Advanced_Engineering\" >Full-Stack Co-Design and Advanced Engineering<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/lockitsoft.com\/?p=7629\/#The_Evolution_of_Inference_Transitioning_to_Agentic_AI_Workloads\" >The Evolution of Inference: Transitioning to Agentic AI Workloads<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/lockitsoft.com\/?p=7629\/#Scalability_and_Rack-Level_Efficiency_of_the_GB300_NVL72\" >Scalability and Rack-Level Efficiency of the GB300 NVL72<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/lockitsoft.com\/?p=7629\/#Software_Velocity_and_Continuous_Post-Submission_Gains\" >Software Velocity and Continuous Post-Submission Gains<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/lockitsoft.com\/?p=7629\/#Democratizing_AI_Infrastructure_Across_the_Ecosystem\" >Democratizing AI Infrastructure Across the Ecosystem<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/lockitsoft.com\/?p=7629\/#Broad_Implications_for_the_Global_Technology_Landscape\" >Broad Implications for the Global Technology Landscape<\/a><\/li><\/ul><\/nav><\/div>\n<h3><span class=\"ez-toc-section\" id=\"Preview_Performance_and_the_Rise_of_the_Vera_Rubin_Architecture\"><\/span>Preview Performance and the Rise of the Vera Rubin Architecture<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Marking a major milestone in high-performance computing, NVIDIA submitted preview results for its next-generation Vera Rubin NVL72 platform during the MLPerf Inference v6.1 evaluation cycle. The submissions focused heavily on two of the industry&#8217;s most computationally demanding benchmarks: the DeepSeek-R1 reasoning model and the Qwen3-VL vision-language model. These workloads represent the cutting edge of AI complexity, demanding exceptional memory bandwidth, sophisticated orchestration, and massive parallel processing capabilities.<\/p>\n<p>The Vera Rubin NVL72 architecture delivered staggering performance gains during testing. Across offline, server, and interactive scenarios on the Qwen3-VL benchmark, the Vera Rubin platform achieved up to 3.7 times higher throughput than the preceding GB300 NVL72 system. This remarkable leap was achieved by leveraging vLLM in conjunction with the open-source NVIDIA Dynamo inference framework. Similarly, on the DeepSeek-R1 benchmark, utilizing the NVIDIA TensorRT-LLM library, Vera Rubin delivered up to 2.5 times higher throughput than its predecessor. These early preview metrics not only highlight NVIDIA\u2019s relentless pace of innovation but also signal how much further performance is expected to climb as continuous software refinements mature ahead of full commercial availability.<\/p>\n<figure class=\"article-inline-figure\"><img decoding=\"async\" src=\"https:\/\/blogs.nvidia.com\/wp-content\/uploads\/2026\/09\/mlperf-inference-6-1.png\" alt=\"NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut\" class=\"article-inline-img\" loading=\"lazy\" \/><\/figure>\n<p>From a financial and operational perspective, these performance figures carry profound implications. Each Vera Rubin NVL72 rack is engineered to output significantly more tokens, service a vastly expanded user base, and generate substantially higher revenue streams than a comparable GB300 NVL72 rack, all while driving down the overall cost per token. <\/p>\n<h3><span class=\"ez-toc-section\" id=\"Full-Stack_Co-Design_and_Advanced_Engineering\"><\/span>Full-Stack Co-Design and Advanced Engineering<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>The extraordinary performance metrics recorded in the MLPerf Inference v6.1 suite are not the result of hardware improvements alone; rather, they are a testament to rigorous full-stack co-design spanning both silicon and software engineering. The Vera Rubin architecture incorporates enhanced Tensor Cores and a redesigned Transformer Engine specifically optimized to accelerate both the prefill and decode stages of the inference lifecycle. Furthermore, the integration of NVFP4 numerical precision significantly reduces the memory footprint across model weights, attention mechanisms, and Key-Value (KV) caches, thereby boosting overall throughput without triggering a noticeable degradation in output quality.<\/p>\n<p>To handle the immense computational load of Mixture-of-Experts (MoE) architectures\u2014which underpin complex models like DeepSeek-R1 and Qwen3-VL\u2014Vera Rubin\u2019s engineering team heavily utilized disaggregated serving techniques. By separating the prefill and decode phases and pairing them with large-scale expert parallelism, the platform achieves unprecedented efficiency. <\/p>\n<p>This hardware-software synergy is anchored by the NVL72 scale-up domain, which relies on sixth-generation NVIDIA NVLink and dedicated NVLink Switches. This networking fabric delivers a phenomenal tenfold increase in packet rates and three times lower latency compared to off-the-shelf Ethernet alternatives. This robust interconnect foundation is what enables advanced parallelism and disaggregation techniques to operate smoothly at a massive rack scale. Proof of this ecosystem readiness was further demonstrated by industry partners such as Nebius, which also submitted Vera Rubin NVL72 preview results and showcased stellar performance metrics.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"The_Evolution_of_Inference_Transitioning_to_Agentic_AI_Workloads\"><\/span>The Evolution of Inference: Transitioning to Agentic AI Workloads<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>As artificial intelligence rapidly transitions from simple text-generation chatbots to sophisticated autonomous agents capable of reasoning, planning, and executing multi-step workflows, the metrics used to measure inference performance must also evolve. Traditional throughput benchmarks, while still essential, fail to capture the multi-faceted nature of agentic AI. <\/p>\n<figure class=\"article-inline-figure\"><img decoding=\"async\" src=\"https:\/\/blogs.nvidia.com\/wp-content\/uploads\/2026\/09\/nvidia-vera-rubin-delivers-up-to-3-7x-better-performance.jpeg\" alt=\"NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut\" class=\"article-inline-img\" loading=\"lazy\" \/><\/figure>\n<p>Addressing this paradigm shift, preview testing on benchmarks such as the SemiAnalysis AgentX framework revealed that the Vera Rubin NVL72 platform delivers up to 30 times better performance than the GB300 NVL72. Recognizing the industry-wide shift toward autonomous workflows, upcoming evaluation suites like the MLPerf Endpoints benchmark are expected to introduce standardized measurements specifically tailored for agentic inference workloads. This ensures that infrastructure evaluations remain aligned with the real-world demands of modern enterprise applications.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Scalability_and_Rack-Level_Efficiency_of_the_GB300_NVL72\"><\/span>Scalability and Rack-Level Efficiency of the GB300 NVL72<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>While preview architectures capture future potential, current deployments rely heavily on proven scalability. Scaling efficiency\u2014defined as how effectively the addition of more graphical processing units translates into proportional throughput gains\u2014remains a primary indicator of infrastructure productivity. NVIDIA achieves high scaling efficiency through a combination of high-bandwidth, low-latency scale-up interconnects within individual racks, high-speed networking between separate racks, and intelligent request orchestration across distributed nodes.<\/p>\n<p>During the MLPerf Inference v6.1 evaluations, NVIDIA\u2019s DeepSeek-R1 submission demonstrated remarkable elasticity, scaling seamlessly from a single GB300 NVL72 rack containing 72 GPUs up to a massive four-rack configuration comprising 288 GPUs. In the offline scenario, the system achieved an astonishing 99% scaling efficiency, meaning that throughput grew almost in direct proportion to the massive influx of added hardware. <\/p>\n<p>Achieving high scaling efficiency is notoriously difficult in distributed computing. Without careful architectural alignment, adding double the hardware often results in diminishing returns due to communication bottlenecks, making infrastructure costs vastly outpace performance returns. The flawless scaling of the GB300 NVL72 proves that hardware, interconnect fabrics, and orchestration software can operate as a unified, cohesive entity. Furthermore, the GB300 NVL72 exhibited exceptional rack-scale efficiency on the WAN 2.2 text-to-video generation benchmark. It achieved 0.65 frames of 720p video per second with an average generation time of 5.7 seconds per video\u2014delivering nine times higher throughput and 7.5 times lower latency than a single isolated compute node.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Software_Velocity_and_Continuous_Post-Submission_Gains\"><\/span>Software Velocity and Continuous Post-Submission Gains<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>NVIDIA&#8217;s development methodology relies heavily on an aggressive cadence of continuous software enhancements, ensuring that hardware assets appreciate in capability long after initial deployment. In the v6.1 benchmark cycle alone, GB300 NVL72 performance on the Qwen3-VL workload improved by up to 1.6 times compared to the previous v6.0 results. These performance leaps were unlocked through optimizations such as lower KV cache precision, expanded kernel fusion, refined execution kernels, and advanced disaggregated serving via vLLM and NVIDIA Dynamo.<\/p>\n<figure class=\"article-inline-figure\"><img decoding=\"async\" src=\"https:\/\/blogs.nvidia.com\/wp-content\/uploads\/2026\/09\/nvidia-gb300-nvl72-scaling-efficiency.jpeg\" alt=\"NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut\" class=\"article-inline-img\" loading=\"lazy\" \/><\/figure>\n<p>Significantly, software engineering efforts did not pause when the MLPerf submission window closed. Post-submission optimizations\u2014though not yet independently verified by MLCommons at the time of publication\u2014have yielded even higher performance metrics on complex workloads like the GPT-OSS-120B language model and DLRMv3 recommendation systems. This continuous software velocity ensures that enterprise customers running NVIDIA infrastructure benefit from compounding performance dividends over the lifecycle of their hardware deployments.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Democratizing_AI_Infrastructure_Across_the_Ecosystem\"><\/span>Democratizing AI Infrastructure Across the Ecosystem<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>The impact of these technological advancements extends far beyond ultra-scale data centers. Demonstrating versatility across the entire computing spectrum, NVIDIA also submitted edge computing results for the Jetson AGX Thor platform. Utilizing NVIDIA TensorRT Edge-LLM on the newly established Edge-Agentic benchmark with the Qwen3.6-27B model, the compact edge processor showcased robust local intelligence capabilities.<\/p>\n<p>Furthermore, the broad participation of NVIDIA\u2019s partner ecosystem in the MLPerf Inference v6.1 round underscores the mainstream enterprise readiness of these platforms. A total of 19 major technology vendors and cloud providers participated, with eight specifically showcasing multi-node Blackwell NVL72 systems. The participating ecosystem includes industry leaders such as ASUS, Microsoft Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, Hewlett Packard Enterprise (HPE), Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure (OCI), Quanta Cloud Technology (QCT), Red Hat, ScitiX, Supermicro, and Wiwynn.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Broad_Implications_for_the_Global_Technology_Landscape\"><\/span>Broad Implications for the Global Technology Landscape<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>The release of the MLPerf Inference v6.1 results, highlighted by the preview dominance of the Vera Rubin architecture and the proven hyper-scalability of the GB300 NVL72, paints a clear picture of the future trajectory of artificial intelligence. As enterprises transition from proof-of-concept AI experiments to fully autonomous, agentic enterprise deployments, the economic pressures surrounding compute efficiency will only intensify. <\/p>\n<p>By aggressively co-designing full-stack solutions\u2014combining specialized silicon innovations like NVFP4 precision and advanced Tensor Cores with high-speed interconnects like sixth-generation NVLink and continuously optimized software frameworks\u2014NVIDIA is setting a rigorous benchmark for the industry. For CIOs and data center architects navigating the capital-intensive world of generative AI infrastructure, these developments offer a vital roadmap: sustainable, revenue-positive AI economics require deep integration across the entire technology stack, capable of scaling effortlessly from the compact edge to the largest AI factories in the world.<\/p>\n<!-- RatingBintangAjaib -->","protected":false},"excerpt":{"rendered":"<p>The economics of artificial intelligence inference are undergoing a fundamental transformation, driven by an intricate interplay of system performance, efficient infrastructure scaling, and continuous software optimization. As enterprises and hyperscalers race to deploy generative AI applications at a global scale, the underlying hardware and software ecosystems must evolve in tandem to keep operational costs manageable &hellip;<\/p>\n","protected":false},"author":3,"featured_media":7628,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[22],"tags":[23,39,25,827,18,24,4298,42,3082,3513,1893,278,1892],"class_list":["post-7629","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-artificial-intelligence","tag-ai","tag-breakthrough","tag-data-science","tag-economics","tag-inference","tag-machine-learning","tag-mlperf","tag-nvidia","tag-preview","tag-results","tag-rubin","tag-unveils","tag-vera"],"_links":{"self":[{"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/posts\/7629","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=7629"}],"version-history":[{"count":0,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/posts\/7629\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/media\/7628"}],"wp:attachment":[{"href":"https:\/\/lockitsoft.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=7629"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=7629"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=7629"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}