DP-BREM: Differentially-Private and Byzantine-Robust Federated Learning with Client Momentum

Xiaolan Gu

34th USENIX Security Symposium (USENIX Security '25) · Day 2 · ML and AI Privacy 1: Federated Learning and Protecting Data

Overview

Federated Learning (FL) has emerged as a crucial paradigm for collaborative machine learning, enabling multiple parties to train a shared model without direct data exchange. While FL inherently offers privacy benefits by keeping raw data local, the model updates themselves can inadvertently leak sensitive information or become vectors for malicious attacks. Xiaolan Gu's presentation, "DP-BREM: Differentially-Private and Byzantine-Robust Federated Learning with Client Momentum," introduces a novel framework designed to tackle these dual challenges of privacy preservation and robustness against adversarial clients in FL environments.

Watch on YouTube · Slides

Visual summary for DP-BREM: Differentially-Private and Byzantine-Robust Federated Learning with Client Momentum by Xiaolan Gu
Visual summary for DP-BREM: Differentially-Private and Byzantine-Robust Federated Learning with Client Momentum by Xiaolan Gu

Key moments

  1. 0:00 Introduction to Federated Learning and its threats
  2. 2:40 Understanding Differential Privacy (DP) and its levels
  3. 5:55 Challenges combining Differential Privacy and Byzantine Robustness
  4. 7:30 Key idea: Learning from History (LFH) with client momentum
  5. 8:50 Optimal placement of DP noise for better utility
  6. 9:40 DP-BREM framework workflow and algorithm steps
  7. 10:40 Extending DP-BREM to DP-BREM+ for a non-trusty server
  8. 11:25 How DP-BREM ensures privacy and robustness

DP-BREM: Differentially-Private and Byzantine-Robust Federated Learning with Client Momentum

Speakers: Xiaolan Gu

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=GGyojPmyDaw

Overview

Federated Learning (FL) has emerged as a crucial paradigm for collaborative machine learning, enabling multiple parties to train a shared model without direct data exchange. While FL inherently offers privacy benefits by keeping raw data local, the model updates themselves can inadvertently leak sensitive information or become vectors for malicious attacks. Xiaolan Gu's presentation, "DP-BREM: Differentially-Private and Byzantine-Robust Federated Learning with Client Momentum," introduces a novel framework designed to tackle these dual challenges of privacy preservation and robustness against adversarial clients in FL environments.

The talk highlights the critical need for FL systems that are simultaneously private and robust, especially in the cross-silo setting where organizations contribute larger datasets. DP-BREM addresses privacy through record-level differential privacy (DP), protecting individual data points within client datasets, and enhances robustness using a momentum-based aggregation method called Learning From History (LFH). This research is significant because it provides a practical and theoretically sound approach to achieving a superior balance between privacy, robustness, and model utility—a notoriously difficult trifecta in distributed machine learning.

Background

▶ Watch: Introduction to Federated Learning and its threats (0:00)

Federated Learning operates on the principle that clients train models locally and send only model updates (like gradients or weights) to a central server, which then aggregates these updates to refine a global model. This model is subsequently redistributed to clients for further local training. FL is typically deployed in two main settings: cross-device FL, involving a large number of users (e.g., smartphone owners) with small datasets, and cross-silo FL, which involves a smaller number of organizations or institutions with more stable and often larger datasets. This paper specifically focuses on the cross-silo scenario.

Despite its inherent privacy advantages over centralized training, FL is susceptible to significant threats:

  1. Privacy Attacks:
  • Membership Inference Attacks (MIA): Attempt to determine if a specific data point was part of the training set.
  • Model Inversion Attacks: Aim to reconstruct original input data by analyzing the trained model.
  • Gradient Leakage: Unique to FL, where attackers on the server side exploit client gradients to recover private training data.
  1. Model Poisoning Attacks (Robustness Threats):
  • Backdoor Attacks: Malicious clients embed hidden triggers into the model, causing specific misclassifications when the trigger is present, while maintaining overall model performance.
  • Data Poisoning Attacks: Malicious clients send arbitrary, manipulative updates to disrupt training, degrade model accuracy, or prevent convergence. The talk specifically focuses on these Byzantine attacks due to their generality and potential for severe model damage.

To mitigate privacy risks, Differential Privacy (DP) is a formal guarantee that quantifies the privacy loss incurred by an algorithm. DP ensures that a model's output does not reveal whether any single individual data point was used during training. This guarantee is controlled by a parameter called epsilon (ε); smaller ε values indicate stronger privacy but typically require adding more noise, which can reduce model accuracy. In FL, DP can be applied at two levels: client-level DP, which protects an entire client's dataset as a whole, and record-level DP, which protects individual data points within each client's dataset. The latter is the focus of DP-BREM, particularly relevant in cross-silo FL where each client might represent an organization with many data points.

Existing approaches to combining DP and robustness in FL face significant challenges:

  • Most existing Byzantine-robust methods rely on median-based aggregators (e.g., Krum, Trimmed Mean) to filter out outliers. However, recent studies have shown that sophisticated, multi-crafted malicious updates can still circumvent these aggregators, slowly degrading the model.
  • Crucially, median-based aggregation is fundamentally incompatible with DP-SGD (Differential Private Stochastic Gradient Descent), which relies on averaging gradients and then adding noise. This mismatch leads to a poor privacy-utility tradeoff when attempting to combine these techniques.
  • Achieving a good balance among privacy, robustness, and utility (model accuracy) is inherently difficult, with prior works often optimizing only two of these dimensions. DP-BREM aims to achieve all three.

Key Findings

▶ Watch: Challenges combining Differential Privacy and Byzantine Robustness (5:55)

The DP-BREM framework introduces several key findings and contributions that significantly advance the state of differentially-private and Byzantine-robust federated learning:

  1. Novel Combination of Momentum-Based Aggregation and Central DP: DP-BREM successfully integrates Learning From History (LFH), a momentum-based aggregation method, with centralized Differential Privacy. This combination is foundational to achieving robust and private FL, as LFH's average-based nature is inherently more compatible with DP-SGD than traditional median-based robust aggregators.
  2. Record-Level Differential Privacy with Bounded Sensitivity: The framework formally guarantees record-level differential privacy by applying gradient clipping at the client level and adding DP noise centrally after aggregating client momentum. A key insight is the analytical proof that even though noise is added after momentum aggregation, the sensitivity of the final aggregated momentum remains bounded, allowing for accurate DP noise calibration.
  3. Additive Error Effects for Improved Convergence: DP-BREM demonstrates that the effects of Byzantine clients and DP noise on model convergence are additive, rather than multiplicative. This crucial finding implies that the presence of adversaries and privacy noise does not amplify each other's negative impacts, leading to significantly better convergence and higher utility (model accuracy) compared to local DP settings where multiplicative effects often hinder performance.
  4. Superior Privacy-Robustness-Utility Tradeoff: Experimental evaluations on datasets like MNIST and FMNIST consistently show that DP-BREM achieves the best trade-off between privacy, robustness, and utility across various attack strategies and privacy budgets (epsilon values). It maintains higher model accuracy even under strong Byzantine attacks and stringent privacy requirements.
  5. Extension for Untrusted Servers (DP-BREM+): The paper extends the core DP-BREM framework to DP-BREM+, which addresses scenarios where the central server is not fully trusted (i.e., honest-but-curious). This is achieved by incorporating secure computation techniques, specifically distributed and jointly generated DP noise using verifiable secure sharing, thereby removing the trust assumption on the server for privacy protection.

Technical Deep Dive

▶ Watch: Optimal placement of DP noise for better utility (8:50)

DP-BREM's design addresses the fundamental challenges of integrating differential privacy and Byzantine robustness in federated learning. The system architecture consists of a central aggregation server and multiple clients. The server orchestrates the training process, while clients perform local model updates. The core assumptions are that the server is honest-but-curious (or fully trusted for the base DP-BREM), and a minority of clients can behave maliciously, sending manipulative updates without affecting honest clients.

The high-level approach of DP-BREM involves three main steps:

  1. Robust Aggregation: Start with a robust aggregator to defend against Byzantine clients.
  2. DP Noise Addition: Add differentially private noise to ensure formal privacy guarantees.
  3. Secure Computation (for untrusted server): Apply secure computation techniques to achieve central DP without requiring a fully trusted server.

The key idea behind DP-BREM's robustness is its reliance on Learning From History (LFH), a momentum-based aggregation method. Unlike traditional aggregators that consider only the current updates, LFH averages a client's updates across multiple training rounds, incorporating a form of "momentum." This approach offers several advantages:

  • Noise Smoothing: By averaging over time, LFH smooths out random noise inherent in individual client updates, making them more stable and reliable.
  • Enhanced Malicious Update Detection: Small, insidious malicious changes that might appear harmless in a single round begin to accumulate and become more noticeable over multiple rounds due to momentum. This accumulation makes Byzantine attacks easier to detect and mitigate compared to methods that only look at instantaneous updates.
  • Compatibility with DP-SGD: Crucially, LFH is an average-based method, making it inherently compatible with DP-SGD, which also relies on averaging. This resolves the fundamental incompatibility issue faced by median-based robust aggregators, paving the way for a more effective combination of privacy and robustness.

The second key design choice is where to add the Differential Privacy noise.

  • Local DP (Client-side noise): If each client adds noise locally before sending updates, it becomes significantly harder for the server to distinguish between honest and malicious clients. This severely compromises robustness, as the noise can mask adversarial behavior.
  • Central DP (Server-side noise after aggregation): Adding noise centrally on the server after aggregating client updates requires substantially less noise to achieve the same privacy guarantee. This significantly improves model utility. While this approach traditionally assumes a trusted server, DP-BREM addresses this by integrating secure computation techniques.

DP-BREM Workflow

The proposed DP-BREM framework operates in rounds:

  1. Server Broadcasts: At the beginning of each round, the central server broadcasts the current global model parameters to all participating clients.
  2. Client-side Computation: Each client receives the global model, computes local gradients based on its private data, and then computes its momentum vector. During this step, record-level gradient clipping is applied to ensure that the contribution of any single data point to the gradient is bounded, which is a prerequisite for differential privacy.
  3. Client Submission: Selected clients send their computed momentum vectors (not raw gradients) to the server.
  4. Server Aggregation and Update: The server aggregates the received momentum vectors from all participating clients. Then, it adds calibrated DP noise to this aggregate. Finally, the server uses this noisy, aggregated momentum to update the global model parameters. The process then repeats for the next round.

Privacy Analysis

The primary challenge in DP-BREM's privacy analysis lies in the interaction between gradient clipping and momentum. Clipping is applied at the individual gradient level on the client side, but the DP noise is added to the aggregated momentum on the server side, which is indirectly derived from these clipped gradients and mixed across runs. The key insight that enables formal privacy guarantees is that the sensitivity of the final aggregated momentum is still bounded. This allows for precise calibration of the DP noise, ensuring record-level differential privacy for individual data points within each client's dataset.

Robustness Analysis

DP-BREM's robustness analysis investigates how model convergence is affected by two primary sources of error: the presence of Byzantine clients and the addition of DP noise. The significant finding here is that the effects of these two error sources are additive. This means that the impact of the noise and the impact of the adversaries do not multiply or amplify each other. Their negative contributions simply add up, leading to better overall convergence properties and higher model utility. This contrasts sharply with local DP settings, where noise and attacks can have multiplicative effects, significantly slowing convergence and degrading performance.

DP-BREM+ for Relaxed Trust Assumptions

For scenarios where the central server cannot be fully trusted with privacy (i.e., an honest-but-curious server that might attempt to infer client data), the framework extends to DP-BREM+. This extension incorporates secure computation techniques to protect privacy even when the server is curious. Specifically, DP-BREM+ utilizes distributed and jointly generated DP noise by clients using verifiable secure sharing. This ensures that the DP noise is correctly generated and added, but no single party (including the server) learns the un-noised aggregate or the individual client contributions, thereby preserving privacy without relying on a fully trusted server. While the talk did not delve into the cryptographic details due to time constraints, the paper provides further information on this critical extension.

Demo / Proof of Concept

▶ Watch: DP-BREM framework workflow and algorithm steps (9:40)

Instead of a live, interactive demonstration, the talk presented compelling experimental results that served as a proof of concept for DP-BREM's effectiveness in achieving both privacy and robustness. The evaluations were conducted on widely used datasets, specifically MNIST and FMNIST (Fashion-MNIST), with similar trends observed on CIFAR.

The experiments focused on two main aspects:

  1. Robustness Evaluation:
  • Setup: The researchers varied the percentage of malicious Byzantine clients in the FL system to simulate increasingly strong attacks. DP-BREM was compared against several strong baselines, all of which incorporated record-level differential privacy and aimed to provide Byzantine robustness.
  • Attack Strategies: The evaluation covered four different attack strategies, with the last one being described as the "most advanced" known to the researchers. These strategies likely involved sophisticated ways for malicious clients to craft their updates to maximize model degradation.
  • Results: DP-BREM consistently demonstrated superior robustness across all tested attacks and datasets. It achieved higher model accuracy even as the percentage of malicious clients increased (i.e., as attacks grew stronger). This indicates that DP-BREM's momentum-based aggregation effectively identifies and mitigates the impact of adversarial updates.
  1. Privacy-Utility Trade-off:
  • Setup: In these experiments, the privacy parameter epsilon (ε) was varied. A smaller ε signifies stronger privacy, but typically necessitates adding more noise, which can reduce model accuracy. The goal was to assess how well DP-BREM maintains utility under different privacy budgets.
  • Results: DP-BREM consistently achieved the best trade-off between privacy and utility, even in the presence of Byzantine attacks. This means that for a given level of privacy (e.g., a specific ε value), DP-BREM maintained higher model accuracy compared to baseline methods. Conversely, for a target model accuracy, DP-BREM could achieve stronger privacy. This highlights the efficiency of central DP noise addition and the additive nature of error sources.

The experimental findings clearly validate the theoretical claims of DP-BREM, demonstrating its practical efficacy in real-world federated learning scenarios where both data privacy and system robustness are paramount.

Defensive Implications

▶ Watch: How DP-BREM ensures privacy and robustness (11:25)

The DP-BREM framework offers several crucial insights and actionable implications for defenders and practitioners deploying or designing federated learning systems:

  1. Embrace Momentum-Based Aggregators: For FL deployments concerned with Byzantine robustness, moving beyond traditional median-based aggregators (like Krum or Trimmed Mean) is essential. Momentum-based methods like Learning From History (LFH), as integrated into DP-BREM, provide superior defense against sophisticated, multi-crafted attacks by leveraging the temporal accumulation of client updates. Their compatibility with averaging also makes them a better fit for differentially private training.
  2. Prioritize Central Differential Privacy: When feasible, implementing centralized differential privacy (where noise is added on the server after aggregation) is significantly more efficient than local DP (client-side noise addition) for achieving a given privacy guarantee. Central DP requires less noise, leading to higher model utility and better convergence. For scenarios where the server is not fully trusted, the DP-BREM+ extension demonstrates that secure computation techniques can effectively bridge this trust gap, enabling the benefits of central DP without a fully trusted server.
  3. Focus on Record-Level DP for Cross-Silo FL: In cross-silo FL settings, where clients are often organizations with large datasets, record-level differential privacy is the appropriate and stronger privacy guarantee. It protects individual data points, which is critical for compliance and stakeholder trust, as opposed to client-level DP which only protects the entire client's dataset.
  4. Understand Additive Error Interactions: Defenders should be aware that the interaction between DP noise and adversarial attacks is not always multiplicative. DP-BREM shows that with careful design (e.g., central DP and robust aggregation), these effects can be additive, leading to better convergence and overall system performance. This understanding can guide the selection of appropriate privacy and robustness mechanisms.
  5. Evaluate Trade-offs Holistically: The research underscores the inherent difficulty in optimizing privacy, robustness, and utility simultaneously. Practitioners should carefully evaluate their specific threat models and privacy requirements to select frameworks that achieve the best possible balance across these three dimensions. DP-BREM provides a strong candidate for scenarios demanding high levels of all three.
  6. Consider Secure Computation for Enhanced Trust: For high-stakes applications or environments with stringent security requirements, the DP-BREM+ extension highlights the importance of incorporating secure computation techniques (e.g., verifiable secure sharing) to remove the trust assumption on the central server regarding privacy. This offers a robust solution for protecting sensitive data even from a curious aggregator.

By adopting the principles and techniques demonstrated by DP-BREM, organizations can build more resilient and trustworthy federated learning systems capable of withstanding both privacy breaches and malicious attacks, thereby unlocking the full potential of collaborative AI.

Key Takeaways

  • DP-BREM is a novel framework for Differentially-Private and Byzantine-Robust Federated Learning that combines client momentum with central DP.
  • It achieves record-level differential privacy for individual data points and strong Byzantine robustness against malicious clients.
  • The framework leverages Learning From History (LFH), a momentum-based aggregator, which is more robust and compatible with DP-SGD than median-based methods.
  • DP-BREM demonstrates that the effects of DP noise and Byzantine attacks are additive, leading to better model convergence and utility compared to local DP settings.
  • Experimental results show that DP-BREM consistently achieves a superior trade-off between privacy, robustness, and model accuracy under various attack scenarios and privacy budgets.
  • The DP-BREM+ extension integrates secure computation techniques to protect privacy even when the central server is only honest-but-curious, removing the need for a fully trusted server.

About the Speaker(s)

Xiaolan Gu is the presenter of the paper "DP-BREM: Differentially-Private and Byzantine-Robust Federated Learning with Client Momentum" at USENIX Security. As the lead presenter, she is a key researcher involved in developing and analyzing this advanced framework for secure and private federated learning. Her work focuses on addressing the critical challenges of privacy preservation and robustness against malicious behavior in distributed machine learning environments.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Legitimate academic contribution to a hard problem — combining differential privacy and Byzantine robustness in federated learning without the usual privacy-utility cliff. The additive error interaction result is the real finding here, and LFH's compatibility with DP-SGD is a genuine insight. But this is a USENIX Security paper presentation, not a security practitioner talk, and it reads like exactly that.

Heather Calloway (CISO) — WEAK

Technically credible research on a real problem in federated learning — but it never climbs out of the lab. The defensive implications section lists six bullet points that read like a research paper's conclusion section, not guidance for anyone who actually deploys FL in production.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)