Lessons Learned From Architecting the Highest-scale Operational Systems in the World - Artur Bergman

Artur Bergman

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In this insightful KubeCon EU talk, Artur Bergman, founder and CTO of Fastly, shared invaluable lessons gleaned from over a decade of architecting and operating one of the world's highest-scale content delivery networks (CDNs). Moving beyond conventional technical deep dives, Bergman presented a holistic philosophy for building robust, reliable, and user-centric platforms. His talk emphasized that operating at Fastly's immense scale—handling tens of millions of requests per second and terabits of traffic—demands a fundamental shift in perspective, particularly regarding platform design, performance metrics, and incident management.

Watch on YouTube

Visual summary for Lessons Learned From Architecting the Highest-scale Operational Systems in the World - Artur Bergman by Artur Bergman
Visual summary for Lessons Learned From Architecting the Highest-scale Operational Systems in the World - Artur Bergman by Artur Bergman

Key moments

  1. 0:00 Introduction to Fastly, its scale, and open source support
  2. 4:00 Defining the unique challenge of building a platform
  3. 4:30 Introducing Festool as a metaphor for platform definition
  4. 5:00 Festool's foundational innovations: dust extractors and straight rails
  5. 6:30 Demonstrating Festool's integrated, compatible, and stackable system
  6. 8:00 Festool's innovative Bluetooth battery system and smart features

Lessons Learned From Architecting the Highest-scale Operational Systems in the World - Artur Bergman

Speakers: Artur Bergman, Founder and CTO, Fastly

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=XelZnqR2t2s

Overview

In this insightful KubeCon EU talk, Artur Bergman, founder and CTO of Fastly, shared invaluable lessons gleaned from over a decade of architecting and operating one of the world's highest-scale content delivery networks (CDNs). Moving beyond conventional technical deep dives, Bergman presented a holistic philosophy for building robust, reliable, and user-centric platforms. His talk emphasized that operating at Fastly's immense scale—handling tens of millions of requests per second and terabits of traffic—demands a fundamental shift in perspective, particularly regarding platform design, performance metrics, and incident management.

Bergman's presentation challenged common industry paradigms, advocating for a deeper understanding of what truly constitutes a "platform" and how to maintain its inherent "promise" to users. He explored critical operational strategies, from an obsessive focus on extreme outliers in performance data to a pragmatic approach to preventing global outages and leveraging emerging technologies like large language models (LLMs) for incident analysis. This article delves into the core tenets of Bergman's philosophy, offering actionable insights for engineers and architects grappling with the complexities of large-scale distributed systems.

Background

▶ Watch: Introduction to Fastly, its scale, and open source support (0:00)

Fastly, founded by Artur Bergman in 2011, began with a modest five Points of Presence (POPs) and a small network. Over 14 years, it has evolved into a global powerhouse, boasting approximately 100 physical server locations worldwide, peaking at 40 million requests per second, and managing over 400 terabits per second of total network capacity. A significant portion of this infrastructure supports open-source projects through Fastly's "Fast Forward" program, serving about one million requests per second for software packages. This journey of exponential growth provided Bergman and his team with unique challenges and opportunities to learn what it truly takes to operate systems at an unprecedented scale.

A central theme of Bergman's talk was the distinction between merely building a scalable system and constructing a platform upon which others can reliably build their own scalable systems. This distinction, he argued, requires a more profound understanding of user expectations and the intrinsic "promise" a platform makes. To illustrate this, Bergman drew an extended analogy with Festool, a 100-year-old German power tool brand renowned for its meticulously integrated ecosystem of tools. Festool's products, from dust extractors (high-end vacuum cleaners) to circular saws and routers that run on straight rails, are designed to work seamlessly together. Each new tool, housed in stackable Sustainers, enhances the overall utility of the system without breaking compatibility or user expectations. The Festool battery system, with its Bluetooth integration for dust extractor control and battery tracking, exemplifies how thoughtful innovation can extend a platform's value. This cohesive design fosters user loyalty, as customers inherently trust that new Festool products will integrate flawlessly, reducing cognitive overhead and accelerating value realization. Bergman posited that a true platform, much like Festool or Lego, maintains a consistent "promise" to its users. Any deviation or "break" from this promise—like Lego introducing incompatible bricks—leads to user frustration and erodes trust. This philosophical foundation underpins Fastly's approach to designing and evolving its own high-scale platform.

Key Findings

▶ Watch: Introducing Festool as a metaphor for platform definition (4:30)

Artur Bergman's talk distilled several critical findings from Fastly's operational experience:

  1. The Platform as a Promise: A successful platform is defined by a consistent, implicit promise to its users. This promise dictates that new features, products, or functionalities must integrate seamlessly and enhance the existing ecosystem without introducing friction or breaking established workflows. Breaking this promise, even for perceived short-term gains, erodes user trust and platform stickiness.
  2. Obsessive Focus on Outliers: At extreme scale (e.g., 40 million requests per second), traditional metrics like medians are misleading. Even the 99.9th percentile (P99.9) represents a substantial number of events (40 times per second at Fastly's peak). Ignoring these "outliers" means ignoring a significant portion of user experience and potential systemic issues. Investigating and fixing performance issues at the P99.9 percentile or higher often leads to improvements across all percentiles, benefiting the entire user base.
  3. Outcome-Oriented Development Over "Technical Debt": Innovation and development efforts should be rigorously aligned with achieving clear user outcomes. Bergman cautioned against the indiscriminate use of "technical debt" as a justification for rewrites or changes, suggesting it often masks a desire to work with new technologies or a dislike of existing code. Instead, technical debt should be framed in terms of its operational value and the cost-benefit of maintaining or replacing it, ensuring changes genuinely improve the system for users.
  4. Global Outages are Caused by Global State Changes: The primary drivers of global outages are global state changes, predominantly through software releases, configuration updates, or sudden workload shifts. Bergman emphasized treating configuration changes with the same rigor as code releases, as both can have equally catastrophic global impacts. Prevention hinges on minimizing immediate global changes, implementing robust canary deployments, and prioritizing rapid, ideally automated, rollback capabilities.
  5. Leveraging AI for Incident Analysis: Large Language Models (LLMs) offer a powerful new tool for incident management. By feeding historical incident data into LLMs, organizations can achieve faster classification, improved recall of past events, and automated summarization of complex investigations. This capability is particularly valuable for analyzing "near incidents"—events that didn't lead to full outages but still represent failures from which to learn, often overlooked due to time constraints.
  6. The Imperative of Sound Abstractions: Clear, well-defined abstractions are the bedrock of maintainable and scalable architectures. Without a coherent plan and proper abstractions, systems can devolve into unmanageable "Winchester Mystery House" scenarios, characterized by illogical connections, dead ends, and a complete lack of architectural coherence, hindering future development and operational stability.

Technical Deep Dive

▶ Watch: Festool's foundational innovations: dust extractors and straight rails (5:00)

Fastly's operational environment is a testament to extreme scale. With approximately 100 physical server locations globally, the network handles a peak of 40 million requests per second and sustains over 400 terabits per second of data transfer. This infrastructure is not purely virtualized; it involves physical servers in data centers, a reality that many modern software developers might not encounter. A substantial part of this capacity is dedicated to the "Fast Forward" program, which serves over 1 million requests per second for open-source software packages, highlighting Fastly's commitment to the open-source community.

A cornerstone of Fastly's operational philosophy is its rigorous approach to performance monitoring and debugging, particularly its focus on outliers. Bergman strongly advocated against looking at medians, instead urging engineers to concentrate on the 99th percentile (P99), P99.9, P99.99, and even P99.999 (six nines). At 40 million requests per second, a P99.9 event still occurs 40 times per second, representing a non-trivial number of user experiences. Bergman provided a specific example where debugging an issue at the P99.999 level revealed an unnecessary file I/O operation occurring within an event loop. Fixing this particular outlier not only dramatically reduced latency for the most affected users (from approximately 80 milliseconds at P99.999 to 5 milliseconds at P99.99) but also yielded a measurable improvement across lower percentiles, demonstrating that addressing extreme cases often optimizes the entire system. Fastly's culture embeds this focus, ensuring that outliers are investigated, not ignored, as they often signal deeper systemic issues that sophisticated users will inevitably encounter and report.

The prevention of global outages is reframed at Fastly as understanding "how to cause a global outage." The primary culprits are identified as global state changes, specifically software releases and configuration releases. Bergman stressed that there is no fundamental difference between pushing new code and flipping a configuration flag that enables dormant code; both are equally potent vectors for system-wide failure. The only truly unavoidable global immediate state change in a network context is BGP routing, which propagates almost instantaneously. For all other changes, Fastly adheres to several critical rules:

  • No fast rules, but fast recovery: Prioritize rapid incident response and recovery over rigid, slow deployment processes.
  • Canary Production Investigation: Any crash in a canary production environment must trigger an immediate investigation, with no exceptions.
  • Rapid, Automated Rollback: The ability to quickly and automatically roll back both code and configuration changes is paramount.
  • Maintain Situational Awareness: During critical changes, especially in multi-component systems (e.g., network switches), operators must continuously monitor the state of all affected components. Bergman recounted an incident where a pop went offline because network engineers failed to verify that earlier switches had fully recovered before proceeding with further updates, highlighting the critical need for vigilance.

A forward-looking technical strategy discussed was the use of Large Language Models (LLMs) for enhancing incident management. Fastly has found significant value in feeding its entire incident history into LLMs for training. This allows for:

  • Faster Classification: Quickly categorizing new incidents based on past patterns.
  • Improved Recollection: Rapidly retrieving context and solutions from previous, similar incidents.
  • Automated Summarization: Generating concise summaries of complex, multi-faceted investigations that might span numerous Zoom calls, chats, and log files. This capability is particularly beneficial for analyzing "near failures"—events that didn't escalate to full outages but still represent valuable learning opportunities. In the air safety world, near incidents are treated as failures, a paradigm Bergman advocates for operations, as these tools facilitate the deep investigation often forgone due to time constraints.

Finally, Bergman underscored the non-negotiable importance of abstractions. He used the analogy of the Winchester Mystery House—a sprawling, planless mansion with doors leading to walls and stairs to ceilings—as a cautionary tale for architectural teams. Without thoughtful planning and well-defined abstractions, software architectures can quickly become similarly chaotic, unmanageable, and impossible to scale or maintain. Getting abstractions right, understanding the layers "a little bit underneath you, a little bit above you," is crucial for long-term system health.

Demo / Proof of Concept

▶ Watch: Demonstrating Festool's integrated, compatible, and stackable system (6:30)

This talk was primarily conceptual and architectural, focusing on lessons learned and philosophical approaches to operating high-scale systems rather than demonstrating a specific tool or proof of concept. Artur Bergman conveyed principles through real-world anecdotes, detailed operational statistics from Fastly, and compelling analogies like the Festool platform and the Winchester Mystery House. No live demo or technical proof of concept was presented.

Defensive Implications

▶ Watch: Festool's innovative Bluetooth battery system and smart features (8:00)

The lessons from Fastly's high-scale operations offer profound implications for security and reliability defenders:

  1. Shift Monitoring Focus to High Percentiles: Defenders must move beyond average or median performance metrics. Security incidents, especially those related to denial of service, resource exhaustion, or anomalous behavior, often manifest as outliers in performance or error rates. Monitoring and alerting on the P99.9 or higher percentiles of latency, error rates, and resource utilization can provide early warnings of attacks or system degradation that might be masked by aggregate data.
  2. Rigor in Configuration Management: Treat configuration changes with the same, if not greater, scrutiny as code deployments. Many security vulnerabilities and outages stem from misconfigurations. Implement robust version control, peer review, staged rollouts, and automated rollback for all configuration changes, recognizing their potential for global impact.
  3. Automated Rapid Rollback as a Core Defensive Strategy: The ability to quickly revert to a known good state is a critical defense against both accidental outages and successful attacks. Invest heavily in automated rollback mechanisms for all deployments (code and config) to minimize the blast radius and recovery time objective (RTO) of any incident.
  4. Embrace "Near-Miss" Analysis: Adopt an "air safety" mindset by treating "near incidents" or security alerts that didn't lead to full compromise as valuable learning opportunities, not successes. Leverage tools like LLMs for incident analysis to systematically classify, summarize, and learn from these events, identifying weaknesses before they are exploited.
  5. Secure by Design through Abstractions: Prioritize well-defined architectural abstractions in system design. Clear abstractions help define trust boundaries, enforce least privilege, and simplify security analysis. A chaotic, "Winchester Mystery House" architecture is inherently harder to secure, audit, and maintain, increasing the attack surface and operational complexity.
  6. Situational Awareness During Critical Changes: Emphasize the importance of situational awareness during any critical change, especially those affecting core infrastructure (e.g., network devices, load balancers, identity systems). Blindly following a script without verifying the state of interdependent components can lead to cascading failures, as illustrated by the switch update outage.
  7. Outcome-Oriented Security: Frame security investments and efforts in terms of their contribution to desired business outcomes. Avoid "security debt" arguments that lack clear justification for improving the system's resilience or protecting critical assets. Prioritize security measures that genuinely reduce risk and enhance the platform's "promise" to its users.

Key Takeaways

  • Platform as a Promise: A successful platform consistently delivers on an implicit promise to its users, ensuring seamless integration and value delivery for new features.
  • Focus on Outliers: At scale, prioritize monitoring and debugging at high percentiles (P99.9 and above) as these reveal critical systemic issues and impact a significant user base.
  • Outcome-Driven Development: Align all development and operational efforts with clear user outcomes, meticulously evaluating the true operational value of addressing "technical debt."
  • Global State Change Control: Treat all global state changes, including both code and configuration releases, as potential causes of outages, and implement robust, staged rollouts with rapid, automated rollback capabilities.
  • AI for Incident Learning: Leverage Large Language Models (LLMs) to analyze historical incident data, improving classification, recollection, and summarization, especially for "near incidents."
  • Architectural Abstractions are Paramount: Invest in clear, well-defined architectural abstractions to prevent chaotic, unmanageable systems and ensure long-term scalability and maintainability.

About the Speaker(s)

Artur Bergman is the founder and CTO of Fastly, a leading edge cloud platform and content delivery network. A veteran in the tech industry, Bergman founded Fastly in 2011 with the mission to "make the internet a better place" by ensuring "all experiences are fast, safe, and engaging." Prior to Fastly, he was deeply involved in the open-source community, notably serving on the Perl 5 core team in the previous century. His extensive experience in open-source development and architecting high-performance systems has profoundly shaped Fastly's infrastructure and operational philosophy, which he generously shared in this KubeCon EU presentation.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Artur Bergman delivers a brutally honest and deeply insightful talk on architecting and operating systems at extreme scale. Moving beyond typical platitudes, he provides actionable lessons from Fastly's experience, emphasizing an obsessive focus on performance outliers, the platform's implicit "promise" to users, and the critical role of robust abstractions. The pragmatic application of LLMs for incident analysis, particularly for "near incidents," and the rigorous approach to global state changes are particularly valuable. This isn't just a war story; it's a masterclass in operational resilience.

Heather Calloway (CISO) — MUST SEE

Artur Bergman's talk is a masterclass in operational excellence at extreme scale, delivering fundamental lessons that every CISO and executive leader needs to grasp. He articulates how a platform's implicit "promise" to its users dictates architectural and operational rigor, emphasizing that trust is built on consistency and reliability. His insights on the dangers of global state changes, the power of focusing on performance outliers, and the strategic use of AI for incident learning provide a clear path for reducing systemic risk and building true institutional resilience.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025