Defending Reddit at Scale

Pratik Lotia (Senior Security Engineer · Reddit), Spencer Koch (Principal Security Engineer · Reddit)

DEF CON 33 · Day 1 · Main Stage

Overview

In "Defending Reddit at Scale," Spencer Koch and Pratik Lotia, veteran security engineers from Reddit, pull back the curtain on the intricate strategies and architectural decisions behind protecting one of the internet's busiest platforms from distributed denial-of-service (DDoS) attacks and other malicious traffic. The talk delves into Reddit's sophisticated, multi-layered approach to rate limiting and traffic management, highlighting the challenges inherent in securing a service that processes 1.3 trillion requests and 175 petabytes of data weekly. This presentation is a crucial resource for security professionals grappling with high-volume web traffic, offering practical insights into signal collection, architectural patterns for defense in depth, and innovative resiliency techniques.

Watch on YouTube

Visual summary for Defending Reddit at Scale by Pratik Lotia, Spencer Koch
Visual summary for Defending Reddit at Scale by Pratik Lotia, Spencer Koch

Key moments

  1. 0:00 Introduction and talk agenda overview
  2. 1:15 Understanding Reddit's massive scale and traffic
  3. 2:06 Reddit's DOS defense evolution: from whack-a-mole to teams
  4. 3:08 Key foundational signals for DOS prevention
  5. 4:28 In-depth look at TLS fingerprints (Jaw3/Jaw4)
  6. 6:20 Exploring request header fingerprinting techniques

Defending Reddit at Scale

Speakers: Pratik Lotia (Senior Security Engineer, Reddit), Spencer Koch (Principal Security Engineer, Reddit)

Conference: DEF CON

YouTube: https://www.youtube.com/watch?v=yGYR-tE0ljw

Overview

In "Defending Reddit at Scale," Spencer Koch and Pratik Lotia, veteran security engineers from Reddit, pull back the curtain on the intricate strategies and architectural decisions behind protecting one of the internet's busiest platforms from distributed denial-of-service (DDoS) attacks and other malicious traffic. The talk delves into Reddit's sophisticated, multi-layered approach to rate limiting and traffic management, highlighting the challenges inherent in securing a service that processes 1.3 trillion requests and 175 petabytes of data weekly. This presentation is a crucial resource for security professionals grappling with high-volume web traffic, offering practical insights into signal collection, architectural patterns for defense in depth, and innovative resiliency techniques.

The core of their discussion revolves around Reddit's journey from reactive "whack-a-mole" DDoS mitigation to a proactive, globally distributed defense strategy. Koch and Lotia meticulously detail how a combination of foundational, TLS, HTTP request header, and behavioral signals are leveraged to identify and neutralize threats at various points in the network stack. They underscore the critical balance between cost-effective edge-based defenses and context-rich application-level controls, providing a blueprint for organizations aiming to build robust, scalable security infrastructures without relying solely on vendor black boxes.

The talk is particularly relevant in an era where web applications face constant, evolving threats. Reddit's experience, from its scrappy startup beginnings to its current global footprint, offers invaluable lessons on constructing durable defenses. By sharing their methodologies for custom logging, observability, and unique resiliency techniques like the "slow lane" for problematic traffic, Koch and Lotia empower defenders to move beyond generic solutions and craft tailored, intelligent systems capable of withstanding attacks at internet scale.

Background

▶ Watch: Introduction and talk agenda overview (0:00)

Reddit's journey in DDoS prevention reflects a common trajectory for rapidly scaling internet companies, moving from ad-hoc responses to a deeply integrated, architectural approach. In 2019, when Spencer Koch joined Reddit, the DDoS prevention efforts were largely characterized as "whack-a-mole," a reactive posture where responses were designed to be "cheapest, easiest" and then awaited the next attack. This approach, while typical for early-stage growth, became unsustainable as Reddit's traffic surged. The platform now handles an astounding 1.3 trillion requests and 175 petabytes of data weekly, with GraphQL endpoints alone processing 275,000 requests per second. Such immense scale necessitates a sophisticated and proactive defense strategy.

A significant shift occurred around 2021, coinciding with major changes to Reddit's API and user interaction models, which demanded a more systematic way to monitor and manage incoming traffic. This period saw the formation of a dedicated "transport team" focused on these challenges. By 2025, Reddit evolved to have both "traffic" and "transport" teams, specifically tasked with defending incoming traffic and managing its global distribution, demonstrating a mature commitment to DDoS mitigation.

Reddit’s infrastructure primarily leverages AWS for its serving side, complemented by Fastly and Cloudflare as Content Delivery Networks (CDNs). This multi-cloud and multi-CDN strategy is fundamental to their resilience, but it also introduces complexity in maintaining consistent security policies and observability across diverse environments. The speakers emphasize that their approach was born out of a necessity to build in-house solutions rather than simply "pay a vendor to make the problem go away," leading to a deeper understanding and more customized control over their defense mechanisms. The underlying problem Reddit addresses is the fundamental challenge of distinguishing legitimate user traffic from malicious automated attacks at extreme scale, without degrading the user experience or incurring prohibitive costs.

Key Findings

▶ Watch: Reddit's DOS defense evolution: from whack-a-mole to teams (2:06)

The talk reveals several key findings and contributions to the field of large-scale web security:

  1. Multi-Layered Signal Collection is Paramount: Reddit demonstrates that effective DDoS defense relies on collecting a diverse array of signals across the OSI stack. Beyond foundational data like GeoIP and User Agent, incorporating TLS fingerprints (Jaw3/Jaw4) and proprietary request header fingerprints provides increasingly granular insights into the client and its behavior, enabling more precise blocking. Advanced signals derived from network request timings, analytic events, and flow order further refine this detection, moving beyond simple volumetric attacks to identify sophisticated behavioral anomalies.
  1. Strategic Placement of Rate Limiting (Edge vs. Application): A critical architectural decision is where to apply rate limits. Reddit advocates for a "defense in depth" approach utilizing both edge rate limiting (via CDNs like Fastly and Cloudflare) for cost-effective, coarse-grained blocking, and application rate limiting (using an internal library with Redis) for context-rich, per-user or per-endpoint controls. This hybrid strategy optimizes cost and effectiveness, acknowledging that no single layer can handle all attack vectors.
  1. Observability Through Custom Logging and Tooling: Standard CDN logs are often insufficient for detailed incident response. Reddit's experience highlights the necessity of customized logging with "synthetic code" and "synthetic reason" values to explicitly track why a request was rate-limited. Storing these enriched logs in BigQuery and developing internal CLI tools for rapid querying is crucial for understanding attack patterns, tuning defenses, and minimizing false positives.
  1. Innovative Resiliency: The "Slow Lane" for Malicious Traffic: A standout technique is the concept of a "slow lane" or "low lane." By identifying traffic from known bad actors (e.g., Tor exit nodes, IPs with poor reputation), Reddit routes these requests to extremely constrained resource pools. If these requests "die," it's "no big deal," as they don't impact legitimate user-facing SLOs. This method effectively quarantines and degrades malicious traffic without expending valuable resources on mitigation for good users.
  1. Architecting for Defense: Proactive architectural decisions significantly enhance defensive capabilities. Splitting domains for different workloads (server rendering, media, chat), intelligent caching strategies, and isolating backend services into specific Kubernetes pods (e.g., for login vs. post details) allows for targeted rate limiting and easier identification of anomalous behavior against specific service endpoints.
  1. Addressing Global Scale Challenges: As Reddit expands globally ("Worldwide Reddit"), consistency in rate limiting across multiple points of presence becomes a complex problem. The talk identifies issues like maintaining consistent routing (e.g., via anycast or circular hashing) and the expense of replicating KV stores globally as critical considerations for any platform operating at a similar scale.

Technical Deep Dive

▶ Watch: Key foundational signals for DOS prevention (3:08)

Reddit's defense strategy is built on a sophisticated stack of signal collection, architectural principles, and rate-limiting implementations, designed for unparalleled scale.

Signal Collection: The Foundation of Intelligent Defense

The ability to accurately classify incoming traffic is paramount. Reddit employs a progressive stack of signals, moving from coarse-grained to highly specific:

  1. Foundational Signals: These are the basic data points available from almost any proxy or CDN:
  • GeoIP+: Beyond raw IP addresses, Reddit enriches GeoIP data using providers like Digital Elements (via Fastly) to identify hosting providers, Tor exit nodes, and other contextual information about the origin IP. While "not super useful" on its own, it forms a base layer.
  • User Agent: This header is crucial for modeling expected traffic patterns. Anomalies, such as an iOS app user agent hitting a web platform URL, immediately flag suspicious activity.
  • Path and Query: Standard URL components provide context on the requested resource.
  1. TLS Fingerprints: Moving up the stack, TLS fingerprints capture characteristics of the client's TLS handshake. Tools like Jaw3 and Jaw4 (referenced in the talk) allow Reddit to identify the operating system, client software, and underlying libraries making the request. This is a powerful signal because it's harder to spoof than a User Agent string. Challenges include GREASE (which adds random data for interoperability testing, increasing noise) and protocol sprawl (e.g., HTTP/3, QUIC) which modify handshake behaviors, requiring robust fingerprinting logic to filter out irrelevant variations. CDNs like Fastly typically provide this functionality.
  1. Request Header Fingerprints: This is a less-publicized but highly effective signal. While specific documentation is scarce, Reddit hopes to open-source their algorithm. They collaborated with Fastly to integrate their logic, resulting in the Fastly-Info-Fingerprint header. Similarly, Cloudflare provides raw request header names. The core idea is that the way an HTTP request packet is constructed—the order, presence, and values of headers—is highly dependent on the client (e.g., Golang, Python Requests, curl). This allows for fingerprinting the specific client library or tool making the call, providing a deeper layer of identification than TLS alone.
  1. Advanced Signals: These signals require more application context and often asynchronous processing:
  • Network Request Timings: By understanding the typical behavior of their single-page applications (SPAs) like New Reddit (initial burst of asset loading, GraphQL calls, then quiet), Reddit can detect non-human behavior. A simple curl request, for instance, won't load JavaScript or follow typical browser-like request sequences, making it easily distinguishable.
  • Analytic Events: The presence or absence of expected JavaScript-fired analytic events can indicate whether a real browser (and human user) is interacting with the page.
  • Cookies and JOTs (JSON Web Tokens): The absence of expected session cookies or JOTs on requests that should naturally have them (e.g., direct access to a post without a preceding embed or origin) can signal a non-human client.
  • Flow Order: Critical user flows, such as the login process (typically three specific API calls in sequence), can be monitored. Requests that deviate from this expected order (e.g., attempting to brute-force a password by hitting only the final login endpoint repeatedly) are strong indicators of malicious activity.

Rate Limiting Architecture: Edge vs. Application

Reddit employs a "defense in depth" strategy by strategically placing rate limits at two key layers:

  1. Edge Rate Limiting (CDN Layer):
  • Cost-Effectiveness: This is the primary driver. Blocking traffic at the CDN (Fastly or Cloudflare) means it never reaches Reddit's origin servers, saving compute and bandwidth costs.
  • Implementation:
  • Fastly: Reddit utilizes Fastly's Edge Rate Limiter, a paid product. It provides two key-value (KV) stores: a penalty box (for blocked entities with a configurable TTL) and a rate counter (an "abacus" incremented for each request). Reddit uses VCL (Fastly's configuration language) for this, avoiding the more expensive "Compute@Edge" product for basic rate limiting.
  • Cloudflare: For Cloudflare, Reddit has largely replicated Fastly's functionality using Cloudflare Workers and KV Workers. This allows them to maintain a multi-CDN approach and swap providers as needed.
  • Limitations: Edge KVs often offer limited observability and control. Reddit notes that Fastly's KV isn't directly inspectable, which can be a challenge for debugging or fine-tuning. This layer is best for "coarse-grained" blocking (e.g., by IP, or IP + TLS fingerprint).
  • Global Scale Challenge: With "Worldwide Reddit" deploying serving stacks globally, ensuring consistent rate limiting across different points of presence (PoPs) becomes complex. If routing isn't consistent (e.g., via anycast or circular hashing), a single attacker could distribute requests across PoPs, making per-PoP counters ineffective. Replicating caches globally is expensive.
  1. Application Rate Limiting (Origin Layer):
  • Context Richness: This layer is used when more specific application context is needed, such as per-user or per-endpoint logic (e.g., "how many posts can a user make in an hour?"). This requires knowing the authenticated user or the specific API endpoint.
  • Implementation: Reddit provides a Baseplate library for its developers, simplifying the implementation of rate limits. Developers can easily integrate Redis as a KV store, requiring "a handful of lines of code." The guidance for developers is to set limits based on "how fast a human should be able to do this given an order of magnitude" (e.g., 2x or 10x human speed).
  • Use Cases: Ideal for OAuth clients (third-party and first-party) where Quality of Service (QoS) dictates different rate limits (e.g., 1000 requests/second for paying clients vs. 10 for free tiers).

Architectural Principles for Defense

Reddit follows several principles to enhance its defensive posture:

  • Domain Splitting: Historically, Reddit has split workloads across different domains (e.g., reddit.com for server rendering, media.reddit.com for media assets, chat.reddit.com for chat). This allows for distinct rate limits and easier identification of anomalous traffic patterns against specific service types.
  • Intelligent Caching: Caching where possible (e.g., static media assets) reduces the load on backend services and makes attacks more expensive. Personalized feeds, however, cannot be cached at the edge.
  • Workload Isolation: Reddit runs on vanilla Kubernetes. They isolate backend services into specific pods (e.g., separate deployments for the login page, post details page, chat iframe rendering, collectively called "shreddit"). This allows for granular rate limiting, understanding behavior, and even swapping out services if one is under attack.

Observability and Tooling

Critical to managing these defenses is robust observability:

  • Custom Logging: CDN logs are often "terrible" in detail. Reddit customizes its logging to include "synthetic code" and "synthetic reason" values, explicitly stating why a request was rate-limited (e.g., "rate limited on TLS fingerprint," "IP + TLS"). This prevents a deluge of generic 429 errors.
  • Unique Request IDs: A UUID for each request (like Cloudflare's ray ID) is essential for tracking issues and debugging. Reddit implemented this for Fastly.
  • Data Storage and Querying: All network traffic logs, enriched with custom details, are streamed to BigQuery. This is cost-effective for Reddit's massive data volume.
  • Internal Tooling: To democratize access to this data, Reddit developed a Go binary CLI tool that allows engineers to query BigQuery directly using simple commands, eliminating the need to interact with the console or write complex SQL. This speeds up incident response and analysis.

Rate Limiting Logic: Band-Pass Filters

The overall rate limiting strategy is conceptualized as a series of "band-pass filters":

  1. Coarse-grained: Start with broad filters like IP address.
  2. Progressively Narrower: Move to IP + TLS fingerprint, then IP + TLS + "pool" (e.g., logged-out, logged-in, OAuth clients), and finally potentially reCAPTCHA for additional signal.

Each step reduces the allowed rate but increases the permutations and complexity. Exemptions (e.g., for Google crawlers) are managed via edge dictionaries in Fastly or KV stores in Cloudflare.

Demo / Proof of Concept

▶ Watch: In-depth look at TLS fingerprints (Jaw3/Jaw4) (4:28)

The talk "Defending Reddit at Scale" did not feature a live demonstration or a traditional proof of concept of an attack and mitigation. Instead, the speakers provided an extensive architectural deep dive into the systems and methodologies Reddit has implemented and refined over several years to achieve robust DDoS protection. The detailed descriptions of their signal collection mechanisms, the strategic placement of rate limits at the edge and application layers, the custom logging and observability tooling, and the innovative resiliency techniques collectively serve as a real-world proof of concept of their effective defense strategies against internet-scale threats. Diagrams illustrating network ingress, Fastly rate limiting logic, and Cloudflare Worker implementations were provided as visual aids to explain their deployed solutions.

Defensive Implications

▶ Watch: Exploring request header fingerprinting techniques (6:20)

The insights shared by Reddit's security team offer actionable strategies for defenders facing similar challenges:

  1. Embrace Multi-Layered Defense: Do not rely on a single defensive mechanism or vendor. Implement defense in depth by strategically placing rate limits and security controls at both the edge (CDN) and application layers. This ensures that even if one layer is bypassed or overwhelmed, subsequent layers can still detect and mitigate threats.
  1. Invest in Diverse Signal Collection: Move beyond basic IP and User Agent analysis. Actively collect and leverage advanced signals like TLS fingerprints (Jaw3/Jaw4), HTTP request header fingerprints, network request timings, analytic events, and flow order anomalies. The more diverse and granular your signals, the more accurately you can distinguish legitimate users from attackers. Consider enriching GeoIP data with reputation services.
  1. Prioritize Custom Observability: Standard CDN logs are often insufficient. Develop custom logging that explicitly records why a request was blocked or rate-limited, using "synthetic codes" and "synthetic reasons." Stream these enriched logs to a scalable data store like BigQuery and build internal CLI tools for rapid querying and analysis. This is critical for understanding attack patterns, tuning defenses, and minimizing false positives.
  1. Implement the "Slow Lane" Strategy: For traffic identified as suspicious or from known bad sources (e.g., Tor exit nodes, IPs with poor reputation), route it to constrained resource pools that can be degraded or fail without impacting core user-facing SLOs. This makes attacks expensive for adversaries while preserving resources for legitimate users.
  1. Make Attacks Expensive: Utilize techniques like tarpitting and response bloat (e.g., Reddit's use of "badger badger badger" in VCL) to increase the cost for attackers in terms of bandwidth, processing, and time, discouraging sustained assaults.
  1. Architect for Security from the Ground Up: Design your application architecture with defense in mind. This includes splitting workloads across different domains, implementing intelligent caching where possible, and isolating critical services into distinct deployment units (e.g., Kubernetes pods). This allows for more targeted defense and easier identification of anomalies.
  1. Plan for Global Scale: If operating a global service, consider the challenges of maintaining consistent rate limiting across multiple points of presence. This includes ensuring consistent routing (e.g., via anycast or circular hashing) and evaluating the cost-benefit of replicating KV stores globally.
  1. Develop Rapid Rollback Mechanisms: Inevitably, legitimate users will be inadvertently blocked. Ensure you have the ability to quickly flush or reset rate limits or add exemptions (e.g., via edge dictionaries) to minimize user impact during incidents.

Key Takeaways

  • Multi-layered defense using both edge (CDN) and application-level rate limiting is crucial for cost-effectiveness and comprehensive protection at scale.
  • Diverse and granular signals, from foundational GeoIP and User Agent to advanced TLS and HTTP header fingerprints, and behavioral analytics, are essential for accurately identifying and mitigating sophisticated attacks.
  • Custom observability with detailed logging, unique request IDs, and internal querying tools (like Reddit's Go binary for BigQuery) is non-negotiable for effective incident response and continuous improvement of defenses.
  • The "slow lane" or "low lane" resiliency technique allows platforms to gracefully degrade service for malicious traffic without impacting legitimate users or core Service Level Objectives (SLOs).
  • Proactive architectural decisions, such as splitting domains, intelligent caching, and workload isolation, significantly enhance the ability to defend against DDoS and other high-volume attacks.
  • Scaling defenses globally introduces challenges related to consistent routing and distributed state management, requiring careful consideration of cache replication and anycast strategies.

About the Speaker(s)

Spencer Koch is a Principal Security Engineer at Reddit, where he has been instrumental in shaping the platform's DDoS prevention and traffic management strategies for over six years. His tenure at Reddit spans a period of significant growth and evolution for the company's security posture, transforming it from a reactive "whack-a-mole" approach to a sophisticated, multi-layered defense system. Koch brings a deep understanding of internet-scale challenges, particularly in balancing cost optimization with robust security measures for massive traffic volumes.

Pratik Lotia is a Senior Security Engineer at Reddit, working alongside Spencer Koch on the security team. He contributes to the development and implementation of the advanced defense mechanisms discussed in the talk, focusing on practical engineering solutions for protecting Reddit's vast user base and infrastructure. Together, Koch and Lotia represent the engineering talent behind Reddit's successful efforts to defend against persistent and evolving online threats.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

A genuinely useful war story from engineers who actually built the thing and are willing to show the seams — VCL snippets, the BigQuery pipeline, the slow-lane trick, header fingerprint collaboration with Fastly. Not groundbreaking research, but it's the kind of honest, transferable operational content that DEF CON's defender track exists for.

Heather Calloway (CISO) — SOLID

A technically credible, practitioner-to-practitioner walkthrough of how Reddit built a mature DDoS defense architecture. Useful for engineers facing similar scale problems, but it stays firmly in the engineering lane — there's no meaningful treatment of governance, risk ownership, or institutional accountability, and the defensive implications don't translate to decision-making at the leadership level.

→ Top-rated talks at DEF CON 33

All talks from DEF CON 33