Scale Smarter Not Harder: How Extending Cluster Autoscaler Saves Mi... Rahul Rangith & Ben Hinthorne
Rahul Rangith, Ben Hinthorne
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this insightful talk from KubeCon EU, Datadog software engineers Ben Hinthorne and Rahul Rangith unveil a sophisticated approach to Kubernetes node autoscaling that has yielded millions in cost savings for their vast infrastructure. The presentation delves into how Datadog, operating Kubernetes across dozens of multi-cloud clusters with tens of thousands of nodes and hundreds of thousands of pods, tackled the complex challenge of optimizing instance type selection. Their solution leverages and extends the native Kubernetes Cluster Autoscaler to dynamically identify and provision the most cost-efficient, performant, and reliable instance types for diverse workloads.

Key moments
- 0:00 Introduction and DataDog's scale
- 1:09 Introducing DataDog's Node Group Set abstraction
- 2:37 Cluster Autoscaler and expander concept
- 3:59 Understanding Cluster Autoscaler expander strategies
- 4:55 Case study: Least waste expander's bin-packing inefficiency
- 6:00 Quantifying bin-packing inefficiency: significant dollar impact
- 6:40 Beyond bin-packing: other instance selection factors
Scale Smarter Not Harder: How Extending Cluster Autoscaler Saves Millions
Speakers: Ben Hinthorne, Software Engineer, Datadog; Rahul Rangith, Software Engineer, Datadog
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=CQ3Wxg4qNaQ
Overview
In this insightful talk from KubeCon EU, Datadog software engineers Ben Hinthorne and Rahul Rangith unveil a sophisticated approach to Kubernetes node autoscaling that has yielded millions in cost savings for their vast infrastructure. The presentation delves into how Datadog, operating Kubernetes across dozens of multi-cloud clusters with tens of thousands of nodes and hundreds of thousands of pods, tackled the complex challenge of optimizing instance type selection. Their solution leverages and extends the native Kubernetes Cluster Autoscaler to dynamically identify and provision the most cost-efficient, performant, and reliable instance types for diverse workloads.
The core of their innovation lies in developing custom tooling—including an instance type adviser, performance benchmarks, and attribute selectors—integrated with the Cluster Autoscaler's gRPC expander. This allows for highly intelligent, automated decision-making that goes far beyond simple bin packing. By abstracting away the intricacies of node infrastructure, Datadog empowers its product teams to focus purely on application development, while the compute autoscaling team ensures optimal resource utilization and significant cost reductions, demonstrating a powerful model for large-scale Kubernetes operations.
The talk highlights that while bin packing is a crucial aspect of cost optimization, a truly effective autoscaling strategy must also account for instance type performance, cloud provider capacity, and specific application requirements. Datadog's journey from recognizing theoretical bin packing opportunities to implementing a fully automated, dynamic system offers a compelling blueprint for other organizations striving for similar efficiencies in their Kubernetes environments.
Background
▶ Watch: Introduction and DataDog's scale (0:00)
Datadog operates a formidable Kubernetes infrastructure, running from scratch in a multi-cloud environment. This setup comprises dozens of clusters, housing tens of thousands of nodes and hundreds of thousands of pods, all working to serve trillions of data points per hour for over 30,000 customers. Managing such a sprawling environment efficiently is a monumental task, and at its heart is Datadog's compute autoscaling team, responsible for ensuring optimal scheduling, scaling efficiency, bin packing, and cost optimization of node infrastructure.
A key abstraction in Datadog's platform is the NodeGroupSet. A NodeGroup is a cloud provider-agnostic representation of an autoscaling group or managed instance group. A NodeGroupSet is a collection of NodeGroups that fall under the same scheduling domain, allowing applications to specify tolerations or node affinities for the set, rather than individual groups. This abstraction simplifies user onboarding, facilitates the management of diverse instance type options, and provides crucial fallback capacity if one instance type runs out.
For node autoscaling, Datadog relies on the Cluster Autoscaler (CA) due to its multi-cloud support, robust operational experience, and the ability to contribute upstream features. When a pod is pending and requires a scale-up, the CA first identifies all eligible node groups via scheduling simulations. If multiple options remain, it uses an expander to select the "best" node group. Built-in expanders include random, least waste (minimizes wasted resources), price (cheapest option), and priority. Expanders can also be stacked to break ties.
Initially, Datadog, with its focus on bin packing efficiency, opted for the least waste expander. However, a critical case study revealed a limitation: if pods arrive in a specific order, the least waste expander could lead to suboptimal bin packing. For example, scaling an 8-core machine for a 6-core pod, followed by a 16-core machine for a 9-core pod, results in 9 wasted cores. Had the 9-core pod arrived first, a single 16-core machine could have accommodated both, wasting only 1 core. This highlighted that while the least waste expander was a good start, it wasn't sufficient for complex, dynamic workloads.
An analysis of their fleet confirmed significant opportunities for improving bin packing, translating into substantial potential cost savings. However, Datadog also recognized that instance type selection extends beyond just bin packing. Factors like performance (network, CPU, memory, storage), cloud provider capacity, and cluster-specific application needs (e.g., network-intensive applications) are equally vital. Given the multi-cloud, multi-cluster environment, a "one expander fits all" solution was clearly inadequate. This led them to the gRPC expander, a powerful feature of the Cluster Autoscaler that allows users to build a custom gRPC service to implement arbitrary logic for node group selection, providing the ultimate flexibility needed for their complex requirements.
Key Findings
▶ Watch: Cluster Autoscaler and expander concept (2:37)
Datadog's journey to optimize Kubernetes node autoscaling revealed several critical insights and delivered substantial results:
- Significant Cost Savings Potential: Initial analysis showed that improving bin packing efficiency across their fleet could translate into millions of dollars in annual cost savings. Through their implemented solutions, Datadog achieved $4 million in potential yearly savings across dozens of clusters within 1.5 months, and a single cluster migration alone resulted in $2 million in savings. One specific cluster saw a cost reduction of over 60% by switching to a more memory-optimized instance type for its workloads.
- Beyond Bin Packing: While crucial, bin packing is not the sole determinant of optimal instance type selection. Performance characteristics (CPU, memory, network, storage), cloud provider capacity, and specific environmental or application requirements (e.g., network-intensive workloads, newer generation CPUs) must also be factored in.
- The Power of the gRPC Expander: The Cluster Autoscaler's gRPC expander proved to be the pivotal component for implementing custom, sophisticated instance type selection logic. It allowed Datadog to integrate external intelligence and dynamic decision-making into the autoscaling process, moving beyond the limitations of built-in expanders.
- Automation is Non-Negotiable for Scale: Manually managing instance type selection across dozens of clusters, with varying workloads, regional capacities, and dynamic scores, is unsustainable. Datadog's success hinged on automating this process through custom controllers.
- Intelligent Node Group Provisioning is Key: Simply creating a node group for every available instance type drastically degrades Cluster Autoscaler performance (P99 duration exceeding 5 minutes). A targeted approach, where node groups are provisioned based on analyzed requirements and optimal types, is essential to maintain CA responsiveness (P99 duration under 10 seconds).
- Unified Decision-Making with gRPC: The gRPC expander, being a generic gRPC service, can be leveraged by multiple internal components (e.g., NodeGroupSet Controller) beyond just the Cluster Autoscaler, creating a centralized source of truth for instance type recommendations across the autoscaling ecosystem.
- Abstraction Empowers Product Teams: By handling the complexities of node infrastructure and instance type optimization automatically, Datadog's platform allows application teams to deploy workloads without needing to worry about underlying hardware, fostering greater agility and focus on product development.
Technical Deep Dive
▶ Watch: Understanding Cluster Autoscaler expander strategies (3:59)
Datadog's strategy for scaling smarter revolves around two primary goals: accurately identifying the best instance types and then effectively scaling those identified types. This required a multi-faceted approach addressing cost, performance, and reliability, integrated into a cohesive autoscaling ecosystem.
Identifying Optimal Instance Types
To identify the best instance types, Datadog bucketed its criteria into three high-level categories: cost, performance, and reliability.
Cost Optimization (Bin Packing)
The challenge of bin packing thousands of unique and dynamic workloads (some leveraging VPA for dynamic resource requests) necessitated a more sophisticated approach than theoretical calculations. Datadog developed the Instance Type Adviser, a component designed to run comprehensive scheduling simulations.
- Inputs: The adviser takes a set of pods (selected via node selectors, e.g., all pods on a specific
NodeGroupSet) and an Instance Catalog. The instance catalog is a comprehensive repository of all possible instance types, detailing their specifications, including cost per hour, CPU capacity, and memory capacity. - Virtual Node Builder: Internally, the adviser uses a Virtual Node Builder to create virtual nodes matching the specifications from the instance catalog.
- Scheduling Simulations: For the actual scheduling simulations, Datadog utilizes the upstream Kubernetes scheduling framework, which allows for accurate prediction of pod placement and resource utilization.
- Output: The Instance Type Adviser ranks instance types based on their bin packing efficiency, specifically targeting a high requested CPU percentage. For example, it might show that a current instance type like
m6a.8x largeachieves 62% requested CPU, whereas migrating tor6a.4x largecould boost this to 95%, leading to significant cost reductions. The results of these simulations are stored as Custom Resource Definitions (CRDs) within the cluster for consumption by other components.
Performance Benchmarking
Beyond raw resource capacity, the actual performance of an instance type significantly impacts the effective cost and application experience. Datadog conducts granular performance benchmarks for:
- Network performance
- CPU performance
- Memory performance
- Storage performance
These benchmarks allow Datadog to weigh cost against true performance. For instance, an instance type B might be 20% more expensive than instance type A, but if it offers 30% more CPU performance, it could ultimately lead to running 30% fewer instances, resulting in overall cost savings when properly autoscaled on CPU. This provides a crucial layer of intelligence for migration decisions.
Reliability and Preferences
Reliability concerns, such as cloud provider capacity, desired size ranges for pods, and specific performance preferences, vary by environment and region. To manage this flexibility, Datadog built a small library for Celigo selectors. These selectors allow users to define rules against instance type attributes (e.g., M6G, R6G, C6G for deep capacity pools; 8-32 CPUs) and environment attributes (e.g., network-intensive types for edge environments, newer generation types for specific clusters, or experimental types for testing). This empowers Datadog to express complex, conditional preferences for instance type selection, ensuring applications always run on suitable and available infrastructure.
Scaling Optimal Instance Types: The Autoscaling Ecosystem
With the tools to identify optimal instance types, the next challenge was to integrate this intelligence into the autoscaling workflow.
Initial Automation and Validation
Datadog's initial approach involved a human operator reviewing the instance analysis results and manually configuring the gRPC expander (via ConfigMaps or CRDs) to guide the Cluster Autoscaler. This validated their core hypothesis:
- In one cluster, analysis showed
m6g.8x large(balanced CPU/memory) was suboptimal for memory-intensive workloads. The recommendation wasr6g.8x large(higher memory/CPU ratio). - Implementing this change resulted in a significant reduction in the number of instances needed and over 60% cost reduction in that single cluster.
This success confirmed the approach but highlighted the need for further automation. Managing instance types across dozens of clusters, each with unique workloads, regional availability differences, and dynamic capacity changes, proved too complex for manual intervention.
The Instance Score Controller
To automate the scoring process, Datadog introduced the Instance Score Controller. This custom Kubernetes controller:
- Watches the instance analysis results (stored as CRDs).
- Dynamically writes these optimal instance type scores into ConfigMaps.
- The gRPC expander reads these ConfigMaps, allowing it to inform the Cluster Autoscaler's decisions without human intervention.
This automation enabled Datadog to expand their migration strategy across their entire fleet, realizing $4 million in potential yearly cost savings within approximately 1.5 months.
The Challenge of Instance Type Availability
A new problem emerged: what if the "best" instance type identified by the analysis simply didn't exist as a NodeGroup in a given cluster? The Cluster Autoscaler can only scale what's available. A naive solution—creating a NodeGroup for every single instance type (potentially hundreds)—was quickly ruled out due to its severe impact on Cluster Autoscaler performance. In one cluster, this approach caused the CA's main loop P99 duration to exceed 5 minutes, making urgent scale-ups unacceptably slow. After cleanup, this was reduced to less than 10 seconds.
The NodeGroupSet Controller
To efficiently provision the necessary NodeGroups without overwhelming the Cluster Autoscaler, Datadog developed the NodeGroupSet Controller. This controller is responsible for:
- Watching
NodeGroupSetdefinitions and their requirements. - Reconciling the actual
NodeGroupsin the cluster to ensure the optimal set is available. - It considers various requirements, such as the most cost-efficient instance type, instance types capable of scheduling large pods (e.g., up to 64 CPUs and 256GB memory), and fallback options (e.g., two fallbacks for each primary instance type) to ensure capacity and reliability.
Crucially, the NodeGroupSet Controller also needs to know which instance types are "best." It achieves this by making a gRPC request to the gRPC expander. This highlights a key architectural benefit: the gRPC expander is a generic service that can be queried by any component, not just the Cluster Autoscaler, allowing for a centralized source of truth for instance type recommendations.
The Closed-Loop Autoscaling Ecosystem
This architecture forms a powerful, closed-loop autoscaling ecosystem:
- Instance Analysis Tools (Adviser, Benchmarks, Selectors) generate optimal instance type recommendations.
- The Instance Score Controller watches these results and populates ConfigMaps.
- The gRPC Expander reads these ConfigMaps, becoming the central authority for optimal instance type scoring.
- The NodeGroupSet Controller queries the gRPC expander to identify required
NodeGroupsand ensures their presence in the cluster. - The Cluster Autoscaler queries the gRPC expander to select the best
NodeGroupto scale up when pods are pending.
This system proved its efficacy in a cluster initially running m6g.8x large instances. When a new application was deployed, the instance analysis identified c6g.4x large as the new optimal type. The NodeGroupSet Controller, querying the gRPC expander, created the c6g.4x large node group. The Cluster Autoscaler then scaled up these new instances, leading to a $2 million reduction in potential yearly cost savings for that cluster.
Demo / Proof of Concept
▶ Watch: Quantifying bin-packing inefficiency: significant dollar impact (6:00)
While the talk did not feature a live, interactive demonstration of the system in action, Ben Hinthorne and Rahul Rangith provided compelling empirical evidence and detailed case studies from Datadog's production environment that served as powerful proofs of concept.
They illustrated the problem and solution with clear examples:
- Theoretical Bin Packing Case Study: A simple scenario with 6-core and 9-core pods demonstrated how the
least wasteexpander could lead to 9 wasted cores, contrasting it with an optimal scenario of 1 wasted core, providing the initial motivation for their work. - Fleet-Wide Bin Packing Analysis: Graphs showing requested CPU percentages across various node group sets and clusters highlighted the "clear opportunity for improvement," translating theoretical waste into concrete dollar impacts.
- Single Cluster Migration Success: A specific example detailed the migration of a cluster from
m6g.8x largetor6g.8x largeinstances. Accompanying graphs clearly showed a significant reduction in the number of instances needed and a cost reduction of over 60% in that cluster, validating their optimal instance type selection methodology. - Automated Cost Savings Across Fleet: A graph tracking "yearly potential cost savings" over a 1.5-month period demonstrated how their automated system, powered by the Instance Score Controller, translated into $4 million in realized savings across dozens of clusters.
- Full Ecosystem in Action: The talk concluded with a detailed example of a cluster transitioning from
m6g.8x largetoc6g.4x largeinstances following a new application deployment. Visualizations showed a spike in node count as the new application scaled, followed by a shift to the more optimalc6g.4x largeinstances, resulting in a $2 million reduction in potential yearly cost savings.
These real-world results and detailed architectural diagrams effectively served as a robust proof of concept, showcasing the practical impact and effectiveness of their extended Cluster Autoscaler solution.
Defensive Implications
▶ Watch: Beyond bin-packing: other instance selection factors (6:40)
The sophisticated autoscaling strategy developed by Datadog offers crucial insights and actionable recommendations for organizations looking to enhance their Kubernetes security posture, operational efficiency, and cost management:
- Prioritize Intelligent Instance Type Selection: Blindly relying on default autoscaling mechanisms can lead to significant resource waste and potential performance bottlenecks. Implement a system that actively identifies the most cost-efficient, performant, and reliable instance types for your specific workloads.
- Leverage the gRPC Expander: For large-scale or complex Kubernetes environments, the Cluster Autoscaler's gRPC expander is an indispensable tool. It allows for the integration of custom business logic, external data sources (like performance benchmarks, capacity pools), and dynamic decision-making that goes beyond the capabilities of built-in expanders. This flexibility is key to adapting to evolving infrastructure needs.
- Build Comprehensive Instance Analysis Tools: Develop or adopt components like Datadog's Instance Type Adviser to run scheduling simulations. This provides empirical data on bin packing efficiency, translating resource utilization into tangible cost impacts. Supplement this with robust performance benchmarking across various dimensions (CPU, memory, network, storage) to understand the true capabilities and effective cost of different instance types.
- Automate Instance Type Management: Manual intervention for instance type selection and node group provisioning does not scale. Implement custom controllers (like Datadog's Instance Score Controller and NodeGroupSet Controller) to automate the identification, scoring, and provisioning of optimal
NodeGroups. This ensures consistent application of best practices, reduces human error, and allows for rapid adaptation to changing conditions (e.g., capacity fluctuations, new application deployments). - Optimize Node Group Provisioning: Avoid the pitfall of creating excessive
NodeGroupsfor every possible instance type, as this severely degrades Cluster Autoscaler performance. Instead, provisionNodeGroupsstrategically based on actual workload requirements, cost efficiency, and necessary fallbacks. This ensures responsiveness and operational stability. - Embrace Abstraction for Application Teams: By abstracting away the complexities of node infrastructure and instance type selection, platform teams can empower application developers to focus solely on their product. This reduces cognitive load, accelerates development cycles, and allows the platform to dynamically optimize resource allocation without requiring application-level changes.
- Monitor and Re-evaluate Regularly: The "best" instance type is not static. Workloads change, cloud provider offerings evolve, and capacity fluctuates. Continuously monitor bin packing efficiency, cost metrics, and application performance. Regularly re-run instance analysis and adapt your automated provisioning and scaling strategies to maintain optimal efficiency.
- Consider Cross-Cluster Optimization: Explore opportunities for optimizing application placement across clusters, as Datadog plans to do. Preemptively deciding where new applications should go, or migrating existing ones, based on bin packing potential can further enhance overall fleet efficiency and cost savings.
Key Takeaways
- Massive Cost Savings Through Intelligent Autoscaling: Implementing a sophisticated, data-driven approach to Kubernetes node autoscaling can lead to millions of dollars in annual cost savings by optimizing instance type selection and bin packing efficiency.
- gRPC Expander is a Game-Changer: The Cluster Autoscaler's gRPC expander is essential for custom, intelligent scaling decisions, allowing organizations to integrate complex logic and external data sources beyond basic expander capabilities.
- Holistic Instance Type Selection: Optimal instance type selection must balance three key criteria: cost (via bin packing simulations), performance (through granular benchmarks), and reliability (considering cloud provider capacity, size ranges, and environmental preferences).
- Automation is Crucial for Scale: Managing dynamic instance types across large, multi-cluster fleets requires robust automation via custom Kubernetes controllers (e.g., Instance Score Controller, NodeGroupSet Controller) to eliminate manual toil and ensure rapid adaptation.
- Strategic Node Group Provisioning: Avoid over-provisioning node groups, which can degrade Cluster Autoscaler performance. Instead, provision node groups intelligently based on identified optimal types and workload requirements, ensuring both efficiency and availability.
- Abstract Infrastructure for Developer Empowerment: By handling low-level infrastructure details and instance type optimization automatically, platform teams enable application developers to focus on product innovation, while the underlying infrastructure dynamically adapts to meet workload needs.
About the Speaker(s)
Ben Hinthorne and Rahul Rangith are both Software Engineers at Datadog. They are key members of Datadog's compute autoscaling team, which is responsible for managing the node infrastructure for Datadog's vast Kubernetes fleet. Their work focuses on critical areas such as scheduling, scaling efficiency, bin packing, and cost optimizations across dozens of clusters and tens of thousands of nodes. Their expertise lies in developing and implementing advanced solutions to ensure Datadog's applications run on the most efficient and reliable infrastructure, enabling product teams to concentrate on their core development.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk from Datadog engineers presents a genuinely impressive deep-dive into how they've extended Kubernetes Cluster Autoscaler to achieve millions in cost savings. They've built a sophisticated ecosystem of custom tooling, including an instance type adviser, performance benchmarks, and dedicated controllers, all leveraging the CA's gRPC expander. It's a prime example of real-world, large-scale infrastructure optimization, demonstrating significant technical skill and tangible financial impact.
Heather Calloway (CISO) — STRONG ACCEPT
This KubeCon talk from Datadog engineers provides a compelling blueprint for how sophisticated platform engineering can translate directly into substantial business value. By extending Kubernetes Cluster Autoscaler with custom intelligence and automation, Datadog achieved millions in cost savings through optimized instance type selection. While deeply technical, the presentation clearly articulates the operational and financial implications, offering a strong case for investing in intelligent infrastructure governance and automation.