{"id":8043,"date":"2026-09-28T10:29:00","date_gmt":"2026-09-28T10:29:00","guid":{"rendered":"https:\/\/lockitsoft.com\/?p=8043"},"modified":"2026-09-28T10:29:00","modified_gmt":"2026-09-28T10:29:00","slug":"uber-servicescale-controller-enables-multi-orchestrator-kubernetes-workload-management-at-global-scale","status":"publish","type":"post","link":"https:\/\/lockitsoft.com\/?p=8043","title":{"rendered":"Uber ServiceScale Controller Enables Multi-Orchestrator Kubernetes Workload Management at Global Scale"},"content":{"rendered":"<p>Uber has officially unveiled its new ServiceScale controller, a sophisticated architectural advancement designed to allow multiple orchestrators to safely manage and scale identical Kubernetes workloads. Developed by senior software engineers Egor Grishechko and Srikar Paruchuru, the technology represents a fundamental shift in how the ride-hailing giant manages its vast compute infrastructure. By decoupling scaling intent from execution, Uber has successfully engineered a mechanism to facilitate regional failovers without the historical necessity of maintaining expensive, reserved idle capacity across its global data center footprint.<\/p>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_82_2 counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<span class=\"ez-toc-title-toggle\"><a href=\"#\" class=\"ez-toc-pull-right ez-toc-btn ez-toc-btn-xs ez-toc-btn-default ez-toc-toggle\" aria-label=\"Toggle Table of Content\"><span class=\"ez-toc-js-icon-con\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/span><\/a><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/lockitsoft.com\/?p=8043\/#The_Scale_of_Ubers_Infrastructure_Challenge\" >The Scale of Uber\u2019s Infrastructure Challenge<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/lockitsoft.com\/?p=8043\/#The_Problem_Idle_Capacity_and_Regional_Resilience\" >The Problem: Idle Capacity and Regional Resilience<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/lockitsoft.com\/?p=8043\/#Architectural_Innovation_The_ServiceScale_Controller\" >Architectural Innovation: The ServiceScale Controller<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/lockitsoft.com\/?p=8043\/#Navigating_the_Complexity_of_Multi-Writer_Systems\" >Navigating the Complexity of Multi-Writer Systems<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/lockitsoft.com\/?p=8043\/#Impact_and_Broader_Industry_Implications\" >Impact and Broader Industry Implications<\/a><\/li><\/ul><\/nav><\/div>\n<h3><span class=\"ez-toc-section\" id=\"The_Scale_of_Ubers_Infrastructure_Challenge\"><\/span>The Scale of Uber\u2019s Infrastructure Challenge<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>To understand the necessity of ServiceScale, one must first appreciate the magnitude of Uber\u2019s compute environment. The company\u2019s Container Platform team manages a massive fleet spanning more than 100 compute clusters distributed across multiple cloud providers, including Oracle and Google Cloud. This infrastructure supports approximately 4,000 distinct microservices, running on a foundation of 3 million CPU cores. The system processes an average of 1.5 million pod launches every single day. <\/p>\n<p>For years, the company relied on its internal platform, &quot;Up,&quot; which functions as a high-level federation layer for the Kubernetes fleet. Service owners interact with Up to manage builds and define scaling expectations, while the Uber Deployment Controller (UDC) acts as the bridge, reconciling these high-level intentions into granular Kubernetes primitives. While this system has served the company well since its migration to Kubernetes, the evolving demands of regional resilience required a more flexible approach to resource allocation.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"The_Problem_Idle_Capacity_and_Regional_Resilience\"><\/span>The Problem: Idle Capacity and Regional Resilience<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Historically, Uber operated its data centers in an active-active configuration. To ensure uninterrupted service during a catastrophic regional outage, the company maintained significant amounts of reserved idle compute capacity in every region. This &quot;buffer&quot; was designed to absorb the sudden surge in traffic that would occur if a surviving region had to take on the workload of an offline one.<\/p>\n<p>However, as the scale of Uber&#8217;s services grew, the cost of keeping millions of cores sitting idle became economically unsustainable. The engineering team sought a more efficient alternative: repurposing capacity from lower-priority, non-critical workloads during an emergency. By scaling down these low-tier services during a failover event, the company could theoretically free up enough compute power to scale up its high-tier, mission-critical services.<\/p>\n<p>This objective introduced a significant architectural conflict. Both the standard &quot;Up&quot; platform and a new, specialized failover orchestrator now required the ability to dictate scaling decisions for the same workloads. Initially, the team considered expanding the existing UDC to incorporate failover logic. However, they ultimately rejected this path, fearing that adding such high-stakes complexity to the primary deployment controller would introduce significant risk. As Grishechko and Paruchuru noted, a regression within the failover logic could potentially cascade into the standard deployment workflows, destabilizing the entire fleet.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Architectural_Innovation_The_ServiceScale_Controller\"><\/span>Architectural Innovation: The ServiceScale Controller<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>The solution arrived in the form of a custom resource definition (CRD) named &quot;ServiceScale&quot; and its corresponding controller, the Service Scale Controller (SSC). By introducing this layer, Uber allowed multiple independent orchestrators to express their scaling desires through the ServiceScale object. The SSC then performs the heavy lifting, reconciling these competing intents into a unified command for Kubernetes.<\/p>\n<figure class=\"article-inline-figure\"><img decoding=\"async\" src=\"https:\/\/res.infoq.com\/news\/2026\/09\/uber-kubernetes-scaling\/en\/headerimage\/generatedHeaderImage-1790360989750.jpg\" alt=\"Uber Separates Scaling Intent From Execution on Kubernetes Platform\" class=\"article-inline-img\" loading=\"lazy\" \/><\/figure>\n<p>The team intentionally avoided the &quot;easy&quot; path of building an external database or a dedicated coordination service. They recognized that adding another moving part would increase the surface area for failure and make debugging significantly more difficult under the high-pressure environment of an incident. By keeping the scaling intent materialized directly within Kubernetes, engineers gained the ability to inspect the system\u2019s state using standard tools. If a discrepancy arose, operators could simply query the ServiceScale object to see exactly which orchestrator was requesting which resource level. Furthermore, this approach simplified the &quot;failback&quot; process, as both steady-state and temporary failover parameters are persisted within the CRD spec, eliminating the need to reconstruct state from fragmented logs.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Navigating_the_Complexity_of_Multi-Writer_Systems\"><\/span>Navigating the Complexity of Multi-Writer Systems<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>The deployment of the ServiceScale controller was not without technical hurdles. The most significant challenge involved the limitations of Kubernetes informer caches, which can occasionally lag behind the actual state of the cluster. In a typical controller environment, this latency of a few seconds can be catastrophic if the controller relies on status fields to trigger irreversible workflows.<\/p>\n<p>To address this, Uber implemented a &quot;read-your-own-write&quot; consistency guardrail. When the controller performs an update, it attaches its current generation as an annotation and refuses to proceed with subsequent actions until it verifies that its cached data reflects that specific generation. This solution mirrors the staleness mitigation features introduced in Kubernetes v1.36, released in April 2026, which aims to provide native support for these consistency patterns. Industry analysts, such as software engineer Prasad M K, have characterized this as an &quot;API contract problem&quot; rather than a mere cache defect, noting that robust systems must treat version tokens as a fundamental requirement for validation.<\/p>\n<p>A secondary challenge involved the consistency of ReplicaSets. When multiple controllers\u2014the UDC and the SSC\u2014attempted to modify the same resource simultaneously, it resulted in metadata drift, where the spec and the metadata fell out of sync. This drift effectively broke proportional scaling for rolling updates, causing some workloads to hang. Uber responded by implementing fleet-wide observability tools to detect such drift in real-time, accompanied by an automated &quot;healer&quot; function within the UDC designed to patch inconsistent ReplicaSets on the fly.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Impact_and_Broader_Industry_Implications\"><\/span>Impact and Broader Industry Implications<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>The implementation of this architecture has yielded measurable results. According to an academic paper published on arXiv in January 2026 regarding Uber&#8217;s Unified Failover Architecture, the company successfully reduced its steady-state provisioning requirements from 2x capacity to just 1.3x. This efficiency gain allowed Uber to eliminate over one million CPU cores from its provisioned inventory, representing a massive reduction in both operational costs and environmental impact.<\/p>\n<p>The rollout of ServiceScale was a meticulous, year-long process. The team relied heavily on staging environments and canary deployments, utilizing the &quot;kind&quot; testing tool to simulate complex, multi-controller interactions. By the time the rollout reached full maturity, it supported both native Kubernetes Deployments and OpenKruise CloneSets, completing the transition without a single customer-impacting outage.<\/p>\n<p>The implications of Uber\u2019s work extend well beyond its own data centers. As the industry moves toward more complex, multi-orchestrator environments, the challenges of state synchronization and write-latency are becoming central themes in platform engineering. Projects like the CNCF-graduated &quot;Karmada&quot; are currently exploring different approaches to multi-cluster orchestration and failover. <\/p>\n<p>The core lesson shared by the Uber team remains a sobering reminder for distributed systems architects: &quot;Multi-orchestrator systems aren&#8217;t hard because of the APIs. They&#8217;re hard because of everything that happens between writes.&quot; As Kubernetes continues to evolve into the standard substrate for global-scale compute, the techniques developed by Uber for managing multi-orchestrator intent will likely become a blueprint for other large-scale enterprises grappling with similar challenges in resilience and resource efficiency. The shift toward making scaling intent a first-class citizen in the Kubernetes API, rather than a hidden internal state, represents a significant step forward in the maturation of cloud-native infrastructure management.<\/p>\n<!-- RatingBintangAjaib -->","protected":false},"excerpt":{"rendered":"<p>Uber has officially unveiled its new ServiceScale controller, a sophisticated architectural advancement designed to allow multiple orchestrators to safely manage and scale identical Kubernetes workloads. Developed by senior software engineers Egor Grishechko and Srikar Paruchuru, the technology represents a fundamental shift in how the ride-hailing giant manages its vast compute infrastructure. By decoupling scaling intent &hellip;<\/p>\n","protected":false},"author":24,"featured_media":8042,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[136],"tags":[138,4615,1616,293,466,95,735,4448,139,35,4614,137,4255,3696],"class_list":["post-8043","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-software-development","tag-coding","tag-controller","tag-enables","tag-global","tag-kubernetes","tag-management","tag-multi","tag-orchestrator","tag-programming","tag-scale","tag-servicescale","tag-software","tag-uber","tag-workload"],"_links":{"self":[{"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/posts\/8043","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/users\/24"}],"replies":[{"embeddable":true,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=8043"}],"version-history":[{"count":0,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/posts\/8043\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=\/wp\/v2\/media\/8042"}],"wp:attachment":[{"href":"https:\/\/lockitsoft.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=8043"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=8043"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/lockitsoft.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=8043"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}