More Data Please: Hands on Green Cloud Experiments - Leonard Pahlke & Antonio Di Turi
Leonard Pahlke, Antonio Di Turi
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In "More Data Please: Hands on Green Cloud Experiments," Leonard Pahlke and Antonio Di Turi from KubeCon EU delve into the often-overlooked environmental impact of cloud computing. The talk addresses the growing energy consumption of cloud infrastructure and the increasing abstraction that distances developers and operators from the underlying hardware and its power demands. The speakers highlight that while cloud technologies like Kubernetes offer unparalleled scalability, agility, and abstraction, they also come with significant hidden costs, particularly in terms of energy consumption and a diminishing understanding of infrastructure at its deepest layers.

Key moments
- 0:00 Cloud benefits vs. hidden energy costs and software bloat
- 2:25 Shifting focus: bottom-up approach to measure cloud energy
- 4:40 Building custom bare metal device for green cloud experiments
- 6:00 Detailing the K3S cluster setup on NixOS for experiments
- 7:40 Introducing Scaphandra for precise energy consumption measurements
More Data Please: Hands on Green Cloud Experiments
Speakers: Leonard Pahlke, Antonio Di Turi
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=4OdYWliYpPg
Overview
In "More Data Please: Hands on Green Cloud Experiments," Leonard Pahlke and Antonio Di Turi from KubeCon EU delve into the often-overlooked environmental impact of cloud computing. The talk addresses the growing energy consumption of cloud infrastructure and the increasing abstraction that distances developers and operators from the underlying hardware and its power demands. The speakers highlight that while cloud technologies like Kubernetes offer unparalleled scalability, agility, and abstraction, they also come with significant hidden costs, particularly in terms of energy consumption and a diminishing understanding of infrastructure at its deepest layers.
The core of their presentation is a practical, "bottom-up" approach to measuring and understanding the energy footprint of cloud-native workloads. They describe setting up a bare-metal Kubernetes cluster using commodity hardware and then systematically measuring its power consumption under various conditions, from an idle state to running microservices, stress tests, and security tools. This hands-on experimentation aims to demystify cloud energy usage and provide concrete data points for a more sustainable approach to cloud-native development.
This talk is particularly relevant in an era where environmental sustainability is a critical global concern. By providing a methodology and initial findings for measuring cloud energy consumption at the hardware level, Pahlke and Di Turi empower the community to make more informed decisions about infrastructure design, application optimization, and tool selection. Their work serves as a crucial step towards fostering "green cloud" practices, encouraging a shift from simply provisioning more compute to optimizing existing resources for energy efficiency.
Background
▶ Watch: Cloud benefits vs. hidden energy costs and software bloat (0:00)
The evolution of cloud computing over the past two decades, particularly with the rise of Kubernetes celebrating its 10th anniversary, has fundamentally reshaped software development and deployment. The dream of the cloud—offering on-demand scalability, accelerated development cycles through agility, and simplified complexity through various layers of abstraction—has largely materialized. However, this convenience has obscured significant "hidden costs." As Antonio Di Turi noted, the physical infrastructure, often unseen by developers, consumes substantial energy. The increasing abstraction leads to an "infrastructures detachment," where many, including himself, might rarely encounter a physical server. This detachment contributes to a "loss of technical knowledge" about the deepest stack layers.
Concurrently, software itself has grown increasingly large and resource-intensive. Docker images weighing 2GB are not uncommon, a stark contrast to previous decades. This phenomenon is attributed to "memory bloat," "execution bloat," and an "environmental blind spot." The primary driver behind this lack of optimization is often "abundant compute" – the ease and perceived cheapness of simply adding more nodes or larger instances rather than spending time on optimization. The relentless pursuit of "fast market" entry and rapid feature development, or "feature creep," further exacerbates this, pushing developers to prioritize speed over efficiency. Moreover, the sheer number of layers in modern cloud stacks makes it difficult for individuals to deep dive and understand the full energy implications of their choices.
Recognizing this challenge, the speakers advocated for a "bottom-up" approach to understanding cloud energy consumption. Instead of abstracting away the underlying hardware, they sought to start from the source: bare metal. The goal was to build a foundational understanding of how different components and workloads directly translate into energy usage, information that could then be extrapolated to more complex, production-grade cloud environments. This approach directly confronts the problem of high-level, abstracted discussions around sustainability by grounding them in tangible, measurable hardware interactions.
Key Findings
▶ Watch: Shifting focus: bottom-up approach to measure cloud energy (2:25)
The experiments conducted by Pahlke and Di Turi yielded several crucial insights into the real-world energy consumption of cloud-native environments, particularly when viewed from a bare-metal perspective.
Firstly, the idle power consumption of their "cold" Kubernetes cluster (running only K3S) was surprisingly low. While initial expectations for the entire setup (four nodes, switch, networking components) were around 40-100 watts based on plug measurements, the software-based tool Scaphandre reported values closer to 1-5 watts for the sum of the four nodes, with individual nodes consuming less than 1 watt. This significant discrepancy highlighted a critical challenge: software-based tools, while valuable, often do not capture the entire power draw of the system, missing what the speakers termed "power idle" or hardware-related power metrics.
Secondly, the experiments demonstrated that deploying and stressing workloads incrementally increased the cluster's power consumption. Adding the Google Cloud Microservices Demo and then applying stress with stress-ng showed a measurable, albeit initially small, increase in power. Further adding Falco for continuous security monitoring also led to additional power spikes. This confirmed that application complexity and workload intensity directly correlate with increased energy demand, even if the absolute values captured by software were lower than anticipated.
Thirdly, and perhaps most importantly, the speakers discovered that by accounting for the baseline offset (the difference between the plug measurement and the software measurement of the "cold" cluster), the software-based energy metrics from Scaphandre became remarkably accurate for measuring relative changes in power consumption. Once this "always-on" hardware power was added to the software readings, the combined figures closely matched the expected power draw measured at the plug for various workload scenarios. This finding is pivotal for defenders and developers, as it suggests that while software tools might not capture total energy, they can provide reliable insights into the incremental energy cost of specific applications and configurations, provided a proper baseline is established.
Finally, the experiments underscored the complexity of understanding "synergistic effects" between different cloud-native projects. The impact of a security tool like Falco on energy consumption, for instance, might vary significantly depending on the specific microservices it monitors. This highlights the need for further research into how combinations of cloud-native tools and configurations influence overall energy footprints. The speakers emphasized that starting from hardware provides a tangible understanding of resource consumption, which is often lost in highly abstracted cloud environments, ultimately fostering a more informed approach to sustainable cloud practices.
Technical Deep Dive
▶ Watch: Building custom bare metal device for green cloud experiments (4:40)
To conduct their "hands-on green cloud experiments," Pahlke and Di Turi meticulously constructed a dedicated, encapsulated bare-metal environment designed for precise energy measurement. This setup was a deliberate departure from highly abstracted cloud environments, aiming to reconnect with the underlying hardware.
The experimental setup comprised:
- Hardware: Four identical Intel mini PCs served as compute nodes, connected via a network switch. The entire system included active cooling (fans). This homogeneous setup allowed for consistent comparisons, though the speakers noted future experiments could involve diverse architectures like ARM chips.
- Network Topology: One of the four nodes (Node 1) was designated as the gatekeeper, responsible for DNS and NAT. This design ensured that all internet-bound traffic from the other three Kubernetes cluster nodes (Node 2, Node 3, Node 4) routed through Node 1, aiding in environment isolation and measurement.
- Kubernetes Distribution: They opted for K3S, a lightweight Kubernetes distribution, due to its simplicity and ease of setup, particularly within a NixOS environment. The K3S components were deployed with their default configurations, allowing for future experiments to explore the impact of changing elements like Container Network Interfaces (CNIs).
A cornerstone of their reproducible environment was NixOS, a declarative Linux distribution and package manager. NixOS allowed them to define the entire system configuration—including package installations, network settings, and open ports—in a declarative file. This provided:
- Reproducibility: Using Nix flakes, they could guarantee that given the same inputs, the system would always produce the same output, ensuring consistency across experiments and updates.
- Simplicity: It facilitated spinning up a bare-metal Kubernetes cluster with minimal manual intervention.
- Declarative Management: System changes were applied by rebuilding and reapplying the NixOS configuration, simplifying management across multiple machines.
For deploying and managing this multi-machine NixOS environment, they utilized Colina, a project that wraps the entire system closure for deployment and updates.
For energy measurement, the team employed Scaphandre, a Rust-based tool designed to collect energy metrics from workloads. The speakers detailed the various options for measuring energy consumption:
- Wire-level Measurement: This involves placing a physical device, such as a smart meter, directly on the power line to measure energy consumption. This method offers the most straightforward and accurate readings.
- Baseboard Management Controller (BMC): Common in data center servers, BMCs are dedicated chips that provide an API to access hardware metrics like heat and power. This method offers good accuracy but was not available in their mini PC setup.
- Software-based Measurement (e.g., RAPL): This approach correlates CPU events with energy usage. Intel's RAPL (Running Average Power Limit) is one such implementation, providing highly accurate estimations. However, as it's software-based and relies on specific CPU events, it may not capture the total energy consumption of all hardware components (e.g., memory) if they are not compatible with the API. This was the primary software method used in their experiments.
The scope of their measurements was intentionally limited to what could be captured within their encapsulated environment using RAPL and other supported sub-components. They acknowledged that broader energy consumption, including monitors, upstream APIs, or SaaS dependencies, falls outside this immediate scope but contributes to the overall environmental footprint.
The experimental workflow, termed a "day journey," involved a manual, step-by-step process:
- Day 0 (Cold Baseline): The K3S cluster was deployed and left running without additional workloads to establish a baseline power consumption.
- Day 1 (Application Deployment): The Google Cloud Microservices Demo (simulating a web shop with approximately 10 services written in different programming languages) was deployed onto the cluster. To generate load, stress-ng was used to simulate around 10 users per minute interacting with the application.
- Day 2 (Security Integration): Falco, a runtime security monitoring tool, was then added to the cluster to observe its impact on power consumption.
Each stage involved manually deploying the workloads, running Scaphandre for continuous 1-minute measurements, and logging the power metrics for subsequent analysis and graphing. This meticulous, controlled approach allowed them to isolate and quantify the energy impact of each added layer of complexity.
Demo / Proof of Concept
▶ Watch: Detailing the K3S cluster setup on NixOS for experiments (6:00)
The talk itself served as a comprehensive demonstration of their "hands-on green cloud experiments," showcasing the methodology, the custom bare-metal setup, and the quantitative results obtained. While there wasn't a live coding demo, the speakers presented a series of compelling graphs and figures derived from their meticulous measurements, effectively serving as a proof of concept for their bottom-up approach to energy monitoring.
The core of the demonstration involved presenting the power consumption data collected at different stages of their "day journey."
- Day 0: Cold Cluster Power: They first revealed the power consumption of their "cold" K3S cluster (Kubernetes running with no user workloads). The audience was challenged to guess the wattage, and the reveal showed Scaphandre reporting values significantly lower than the expected 40+ watts measured at the plug, often closer to 1-5 watts for the entire four-node cluster. This stark difference immediately highlighted the "offset" issue inherent in software-based measurements.
- Day 1: Microservices and Stress: Graphs illustrated the power consumption after deploying the Google Cloud Microservices Demo and subjecting it to load using stress-ng. As expected, the power consumption increased, with visible spikes corresponding to the stress tests. This visually confirmed that application activity directly translates to higher energy usage.
- Day 2: Adding Security (Falco): Further graphs showed the impact of integrating Falco for continuous security monitoring. This addition led to additional increases and spikes in power usage, demonstrating the energy overhead of security tools.
- The Offset Correction: A pivotal moment in the demonstration was when the speakers presented adjusted graphs. By calculating the difference between the total power measured at the plug for the "cold" cluster and the Scaphandre reading for the "cold" cluster, they established a baseline "idle power" offset. When this offset was added to all subsequent Scaphandre measurements, the resulting power curves aligned much more closely with the total power expected at the plug. This effectively validated Scaphandre's ability to accurately measure relative changes in power consumption, provided the initial "idle" offset is accounted for.
This visual evidence served as a powerful proof of concept, demonstrating that while software tools like Scaphandre may not capture the absolute total power consumption (due to unmeasured idle hardware power), they are highly effective at showing the incremental energy cost of different workloads and configurations. The demonstration underscored the importance of understanding this baseline offset for accurate interpretation of energy metrics in cloud-native environments.
Defensive Implications
▶ Watch: Introducing Scaphandra for precise energy consumption measurements (7:40)
The findings from "More Data Please: Hands on Green Cloud Experiments" offer crucial insights for defenders, particularly those responsible for securing and operating cloud-native infrastructure. The talk highlights that security, like other aspects of cloud operations, has an energy footprint that should be considered as part of a holistic sustainability strategy.
- Quantify the Energy Cost of Security Tools: The demonstration explicitly showed that integrating Falco for continuous security monitoring added to the cluster's power consumption. Defenders should not view security tools as "free" in terms of resources. It's imperative to understand and quantify the energy overhead of various security solutions, including runtime protection, network policies, vulnerability scanners, and logging/monitoring agents. This data can inform choices between different security tools or configurations, optimizing for both security efficacy and energy efficiency.
- Optimize Workloads for Efficiency: The concept of "memory bloat" and "execution bloat" directly impacts a system's attack surface and performance, but also its energy consumption. Defenders should advocate for and contribute to practices that promote efficient code and resource utilization. This includes optimizing container images, choosing efficient programming languages, and fine-tuning application configurations. Leaner, more efficient applications are not only more performant and potentially more secure (smaller attack surface) but also consume less energy.
- Understand Your Infrastructure's Baseline: The discovery of a significant "idle power" offset between plug measurements and software-based tools like Scaphandre is critical. Defenders need to establish a comprehensive understanding of their underlying infrastructure's baseline energy consumption, even when "cold." This allows for accurate interpretation of workload-specific energy metrics and helps identify unexpected power draws that could indicate misconfigurations, resource leakage, or even compromise.
- Leverage Hardware-Aware Metrics: While direct BMC access might not be available in public clouds, the principle of understanding hardware-level power consumption remains vital. Defenders should seek out cloud provider metrics or employ tools that can provide the closest possible approximation of hardware-level power draw for their virtualized environments. This information can guide decisions on instance types, regions (considering local energy grids), and scaling strategies.
- Promote Green Infrastructure Choices: The discussion about different hardware (e.g., ARM vs. Intel), operating systems (e.g., stripped-down kernels like unikernels), and Kubernetes components (e.g., CNIs) influencing energy efficiency extends to security. Defenders can contribute to green initiatives by evaluating the energy implications of underlying infrastructure choices. For instance, deploying security components on more energy-efficient ARM-based clusters where feasible could reduce overall consumption.
- Collaborate for Sustainable Security: The speakers highlighted the importance of communicating synergistic effects back to project maintainers. Defenders should engage with security tool vendors and open-source project maintainers to advocate for energy-efficient design and development. By providing real-world data on the energy footprint of security components, the community can collectively work towards more sustainable security solutions that don't disproportionately increase environmental impact.
In essence, "green cloud" principles extend beyond mere operational cost savings to environmental responsibility. By integrating energy awareness into security strategies, defenders can contribute to a more sustainable cloud-native ecosystem without compromising on protection.
Key Takeaways
- Cloud Abstraction Hides Energy Costs: While offering scalability and agility, cloud computing abstracts away the underlying hardware, leading to a disconnect from the significant energy consumption of infrastructure and workloads.
- Bottom-Up Measurement is Crucial: Starting with bare-metal experiments, as demonstrated with K3S on Intel mini PCs, provides critical insights into the real energy footprint of cloud-native components, bridging the gap between high-level abstractions and low-level hardware impact.
- Software Tools Need Baseline Calibration: Tools like Scaphandre (using RAPL) are effective for measuring incremental power changes but often miss a significant "idle power" or baseline offset from the total plug consumption. Accounting for this offset makes software measurements highly accurate for relative analysis.
- Workloads Directly Impact Energy: Deploying microservices, running stress tests with stress-ng, and integrating security tools like Falco measurably increase power consumption, highlighting the energy overhead of application complexity and operational tooling.
- Efficiency Drives Sustainability: Optimizing software for memory and execution, making informed choices about hardware architectures (e.g., ARM vs. Intel), operating systems (NixOS, stripped-down kernels), and Kubernetes components (e.g., CNIs) can significantly reduce energy consumption.
- Synergistic Effects Matter: The combined energy impact of multiple cloud-native components can be complex and non-linear, necessitating further research into how different project combinations affect overall energy footprints.
About the Speaker(s)
Leonard Pahlke and Antonio Di Turi are practitioners deeply engaged in the cloud-native ecosystem, with a particular focus on the intersection of technology and environmental sustainability. Their talk reflects a hands-on approach to understanding the physical impact of digital infrastructure. Antonio Di Turi, as he mentioned, has primarily experienced cloud computing through abstraction, noting that he only recently saw a physical server for the first time. This perspective highlights a common experience among modern developers and underscores the motivation behind their research into the hidden costs of the cloud. Leonard Pahlke appears to have led the technical setup, demonstrating expertise in NixOS, K3S, and energy measurement tools like Scaphandre. Together, they advocate for a more transparent and accountable approach to cloud resource consumption, encouraging the community to delve into the "bottom-up" realities of hardware energy usage to foster a greener cloud future.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
The talk "More Data Please: Hands on Green Cloud Experiments" presents a meticulously crafted, bare-metal experimental setup to quantify the often-hidden energy consumption of cloud-native workloads. By employing K3S on commodity hardware with NixOS for reproducibility, and using tools like Scaphandre, the speakers provide a crucial methodology for calibrating software-based energy metrics against real-world hardware baselines. This "bottom-up" approach demystifies cloud abstraction, offering actionable insights for developers and defenders to optimize workloads, security tools, and infrastructure choices for genuine energy efficiency.
Heather Calloway (CISO) — STRONG ACCEPT
This talk provides a critical, hands-on approach to demystifying the hidden energy costs of cloud-native infrastructure, moving beyond abstraction to tangible measurement. By establishing a practical methodology for quantifying the energy footprint of workloads and security tools, it offers actionable insights for leaders and operators to drive institutional accountability, optimize resource consumption, and inform strategic decisions around sustainability and cloud architecture.