Diffence: Fencing Membership Privacy With Diffusion Models
Yuefeng Peng
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Membership Inference
Overview
In an era where machine learning models are increasingly deployed across sensitive domains, the privacy of training data has become a paramount concern. This talk, "Diffence: Fencing Membership Privacy With Diffusion Models," presented by Yuefeng Peng, introduces a novel defense mechanism against Membership Inference Attacks (MIAs). MIAs pose a significant threat by allowing an adversary to determine whether a specific data sample was part of a model's training dataset, potentially revealing sensitive personal information. Beyond direct privacy breaches, MIAs are known to serve as stepping stones for more sophisticated attacks, such as data extraction against generative models.
Key moments
- 0:00 Introduction to Membership Inference Attacks (MIAs)
- 2:15 Limitations of current MIA defense strategies
- 3:20 Introducing Diffence: A new, retraining-free defense
- 4:00 Diffence: Modifying input samples with diffusion models
- 6:00 Detailed explanation of Diffence's sample selection
- 8:40 Key results: Diffence significantly reduces MIA attack AUC
- 10:00 Diffence improves the privacy and utility trade-off
Diffence: Fencing Membership Privacy With Diffusion Models
Speakers: Yuefeng Peng
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=8CamXhzmniQ
Overview
In an era where machine learning models are increasingly deployed across sensitive domains, the privacy of training data has become a paramount concern. This talk, "Diffence: Fencing Membership Privacy With Diffusion Models," presented by Yuefeng Peng, introduces a novel defense mechanism against Membership Inference Attacks (MIAs). MIAs pose a significant threat by allowing an adversary to determine whether a specific data sample was part of a model's training dataset, potentially revealing sensitive personal information. Beyond direct privacy breaches, MIAs are known to serve as stepping stones for more sophisticated attacks, such as data extraction against generative models.
The core innovation of Diffence lies in its unique approach: rather than modifying the model's training process or its output predictions, it strategically alters the input samples before they are fed into the target model. This "pre-inference phase" defense leverages the power of diffusion models to reconstruct input images, ensuring that the model never encounters exact replicas of its training data. This method is designed to be plug-and-play, not requiring costly retraining of already deployed models, and is compatible with existing privacy-preserving techniques, offering a significant enhancement to the privacy-utility trade-off in machine learning.
This work addresses critical limitations of prior MIA defenses, which typically demand model retraining or additional reference data, often at the cost of utility or significant computational expense. Diffence demonstrates superior performance in reducing MIA success rates across various datasets and model architectures, setting a new state-of-the-art in membership privacy. By offering an effective, lossless, and easily integrable solution, Diffence provides a powerful tool for organizations to bolster the privacy posture of their machine learning systems without disrupting existing infrastructure or compromising model performance.
Background
▶ Watch: Introduction to Membership Inference Attacks (MIAs) (0:00)
The landscape of machine learning security has seen a rapid increase in sophisticated attacks, with Membership Inference Attacks (MIAs) emerging as a particularly insidious threat to data privacy. To understand the context of Diffence, it's essential to first review the standard machine learning pipeline and the existing defense paradigms against MIAs.
A typical machine learning pipeline involves several stages:
- Data Collection: Gathering a dataset of samples for training.
- Model Initialization: Setting up an initial model architecture.
- Model Training: Iteratively updating the model's parameters using the training data and a specific algorithm.
- Final Trained Model: The resulting model, ready for deployment.
- Inference: Using the trained model to make predictions on new, unseen samples.
Existing defenses against MIAs are generally categorized into two main groups, based on where they intervene in this pipeline:
Training Phase Defenses
These defenses aim to reduce the model's memorization of its training data during the training process itself. They typically involve modifying the training algorithms to incorporate privacy-preserving mechanisms. Examples include:
- ATV RG (Adversarial Training for Variational Autoencoders with Regularization Guarantees): Aims to make the model less distinguishable between members and non-members by introducing adversarial examples during training.
- DPSGD (Differentially Private Stochastic Gradient Descent): A foundational technique that adds noise to gradients during training to provide formal differential privacy guarantees. This reduces the ability of an attacker to infer individual training data points.
- Selena: A state-of-the-art training phase defense mentioned in the talk, known for its effectiveness.
While effective, training phase defenses suffer from a crucial limitation: they require retraining the model. For large, complex models or those already deployed in production, retraining can be prohibitively costly, time-consuming, and may not even be feasible if the original training data or infrastructure is no longer available. Some methods also demand additional reference data, which might not be readily accessible in real-world scenarios.
Post-Inference Phase Defenses
These defenses operate after the model has made its predictions but before those predictions are revealed to a potentially malicious party. They modify the model's outputs (e.g., logits, probabilities) to obscure membership information. The goal is to make it difficult for an attacker to differentiate between members and non-members by analyzing the final predictions.
- MemGuard: An example of a post-inference defense that adds noise to the model's output logits to obfuscate membership signals.
Similar to training phase defenses, post-inference methods, while useful, often grapple with trade-offs between privacy enhancement and model utility. Moreover, they still operate on the model's direct output, which may contain residual information about the training data, limiting their ultimate effectiveness.
The limitations of these existing approaches—specifically the necessity for retraining, the requirement for additional data, and the persistent challenge of optimizing the privacy-utility trade-off—highlighted a significant gap in the MIA defense landscape. This gap motivated the development of Diffence, a novel pre-inference defense designed to overcome these hurdles by modifying the input itself.
Key Findings
▶ Watch: Introducing Diffence: A new, retraining-free defense (3:20)
Diffence introduces a paradigm shift in defending against Membership Inference Attacks (MIAs) by operating at the pre-inference phase, offering a suite of key findings that significantly advance the state of the art in privacy preservation for machine learning models:
- No Retraining Required: A cornerstone finding of Diffence is its ability to protect models without demanding costly and time-consuming retraining. Unlike traditional training phase defenses, Diffence is a plug-and-play solution that can be applied directly to already trained and deployed models, making it highly practical for real-world applications.
- Optional Data Requirements: Diffence is effective even without additional reference data. While the availability of some member or non-member samples can further optimize its performance (Scenario 1), the defense remains robust and effective in scenarios where the defender has no prior knowledge of membership (Scenario 3).
- Significant Enhancement of Privacy-Utility Trade-off: The most compelling finding is Diffence's ability to substantially improve the balance between model utility (accuracy) and privacy leakage. Evaluations demonstrate that Diffence pushes models towards an "ideal defense" state, characterized by high utility and low privacy leakage.
- For an undefended model, the average attack AUROC (Area Under the Receiver Operating Characteristic curve, a common metric for MIA success) was reported at 79%.
- When a state-of-the-art training phase defense like Selena was applied, the attack AUROC dropped to 62%.
- Crucially, by applying Diffence on top of the Selena-defended model, the attack AUROC further plummeted to 56%, establishing a new state-of-the-art in MIA defense performance. This reduction from 79% to 56% represents a substantial decrease in an attacker's ability to infer membership.
- Seamless Integration with Existing Defenses: Diffence is designed to be complementary. It can be effortlessly integrated with both training phase defenses (e.g., DPSGD, Selena) and post-inference phase defenses (e.g., MemGuard). This layered approach allows for cumulative privacy benefits, as demonstrated by the improved performance when Diffence is combined with Selena.
- Robustness Across Diverse Settings: The evaluation of Diffence was comprehensive, spanning five different datasets, four distinct model architectures, and assessing its resilience against six state-of-the-art MIA attacks. The results consistently showed Diffence's effectiveness across these varied settings, indicating its generalizability and robustness. The highest attack performance (AUROC) across these six attacks was reported as the measure of privacy leakage, demonstrating that Diffence is resilient even against the most potent adversaries.
- Lossless Utility (When Configured): Through its intelligent sample selection mechanism, Diffence can ensure that there is no model accuracy loss. By selecting a reconstructed sample whose predicted label matches that of the original input, the defense prioritizes preserving the model's utility while enhancing privacy.
- Effectiveness Against Adaptive Attackers: The speaker confirmed during the Q&A that their evaluation assumed an adaptive attacker who was fully aware of the Diffence deployment. Even under this strong adversarial assumption, Diffence maintained its high performance, suggesting its resilience against future attack developments.
In summary, Diffence presents a highly effective, flexible, and practical defense against MIAs, offering significant privacy gains without the prohibitive costs associated with retraining or compromising model utility. Its ability to integrate with and enhance existing defenses positions it as a critical advancement in securing machine learning pipelines.
Technical Deep Dive
▶ Watch: Diffence: Modifying input samples with diffusion models (4:00)
Diffence introduces a fundamentally different approach to membership inference defense by intervening at the pre-inference phase. Instead of altering the model's internal workings or its output, Diffence modifies the input data itself before it ever reaches the target model. This section delves into the technical specifics of how Diffence achieves this.
The core intuition behind Diffence is elegantly simple yet powerful: if the target model never encounters an input that is an exact replica of a sample it saw during training, it becomes significantly harder for an attacker to differentiate between a "member" (a sample from the training set) and a "non-member" (a sample not from the training set). Diffence accomplishes this by using a diffusion model to reconstruct the original images.
The Diffusion Model at the Core
At the heart of Diffence is a pre-trained diffusion model. Diffusion models are a class of generative models capable of generating highly realistic images by iteratively denoising a random input. In Diffence, when an input sample (e.g., an image) is provided, it is first passed through this diffusion model. The diffusion model doesn't simply pass the image through; it effectively "reconstructs" it. This reconstruction process introduces subtle, imperceptible perturbations or alterations to the image, creating a new version that is semantically similar to the original but numerically distinct. This reconstructed image is then sent to the target classification model for inference.
The key benefit here is that the target model, which was trained on the original, unmodified dataset, now receives a slightly altered version of any input. This subtle change is enough to blur the lines between members and non-members from the model's perspective, making it difficult for an MIA to exploit the model's memorization of specific training examples.
Stochastic Generation and Informed Selection
A crucial aspect of diffusion models is their stochastic nature in sample generation. When a diffusion model reconstructs an image, it doesn't produce a single, deterministic output. Instead, due to the inherent randomness in the denoising process, it can generate multiple slightly different versions of the same input. Diffence leverages this stochasticity by generating multiple reconstructed images for each original input sample.
This multi-sample generation is critical for two reasons:
- Mitigating Accuracy Degradation: The inherent randomness can sometimes lead to reconstructed images that, while visually similar, might cause the target model to misclassify them. Generating multiple samples increases the probability of finding a high-quality reconstruction that preserves the original classification.
- Informed Selection for Enhanced Defense: By having a pool of reconstructed samples, Diffence can perform an "informed selection" to pick the best candidate, not just for utility preservation but also for maximizing privacy.
The selection process is guided by two primary criteria:
- Preservation of Model Accuracy: The first and most critical criterion is to select a reconstructed sample whose predicted label by the target model matches that of the original image. This ensures that Diffence operates in a "lossless" manner regarding the model's primary task performance (classification accuracy). The speaker noted that "almost in all the cases there's always a sample that has the same label as original sample," indicating the feasibility of this criterion.
- Normalization of Predicted Logits: Beyond label matching, Diffence aims to further reduce the discriminability between member and non-member predictions. It achieves this by selecting a reconstructed sample whose predicted logit (the raw output of the model before the softmax function) falls within a precomputed range.
- By forcing the logits of all processed samples (both members and non-members) to conform to a specific range, Diffence minimizes the statistical differences in the model's output that MIAs typically exploit. This effectively "normalizes" the output space from a privacy perspective.
Precomputed Logit Range Scenarios
The definition of this precomputed logit range depends on the defender's knowledge about the membership of certain samples. Diffence outlines three scenarios:
- Scenario 1 (Known Members and Non-Members): In this ideal scenario, the defender has access to a small set of both member and non-member samples. These reference samples are used to empirically determine an optimal logit range that effectively blurs the distinction between the two groups. This scenario yields the most effective defense.
- Scenario 2 (Known Members Only): The defender only has access to some member samples. While less information is available, these can still be used to define a range that helps in reducing privacy leakage.
- Scenario 3 (No Membership Knowledge): This is the most challenging scenario, where the defender has no prior knowledge of any sample's membership status. In this case, the range must be determined heuristically or based on statistical properties of the model's outputs on a general dataset. Despite this limitation, Diffence remains effective, demonstrating its practical applicability even in restrictive environments.
Plug-and-Play Integration
A significant technical advantage of Diffence is its plug-and-play nature. It operates as an independent module at the pre-inference stage, meaning it can be seamlessly integrated into existing machine learning pipelines without requiring any changes to the target model itself. This architectural flexibility allows for direct application to models already in production.
Furthermore, Diffence can be combined with other existing MIA defenses:
- Training Phase Integration: A model can first be trained using privacy-preserving algorithms like DPSGD or Selena. Then, during inference, Diffence can be deployed to further enhance the privacy and utility trade-off.
- Post-Inference Phase Integration: Diffence can also work in conjunction with post-inference defenses like MemGuard, creating a multi-layered defense strategy.
By intelligently altering inputs with diffusion models and carefully selecting the best reconstructed samples, Diffence provides a robust, flexible, and highly effective mechanism to fence membership privacy in machine learning systems.
Demo / Proof of Concept
▶ Watch: Key results: Diffence significantly reduces MIA attack AUC (8:40)
The presentation of "Diffence: Fencing Membership Privacy With Diffusion Models" primarily focused on elucidating the methodology, theoretical underpinnings, and comprehensive experimental evaluation of the proposed defense mechanism. While the talk provided detailed results through tables and plots illustrating the performance improvements, it did not include a live demonstration or a step-by-step walkthrough of a practical proof-of-concept during the conference session.
Instead of a live demo, the efficacy of Diffence was thoroughly validated through extensive empirical studies. The speaker presented quantitative data showcasing the reduction in Membership Inference Attack (MIA) success rates (measured by AUROC) across diverse datasets, model architectures, and adversarial attack strategies. The results, such as the drop in attack AUROC from 79% (undefended) to 56% (Diffence combined with Selena), served as the primary evidence of the defense's practical utility and effectiveness. The availability of the paper and code online, as mentioned by the speaker, implies that researchers and practitioners can reproduce and experiment with the Diffence implementation themselves, effectively serving as an accessible proof-of-concept for those interested in its technical details.
Defensive Implications
▶ Watch: Diffence improves the privacy and utility trade-off (10:00)
The introduction of Diffence brings several profound implications for defenders grappling with the challenge of membership privacy in machine learning systems:
- Immediate Protection for Deployed Models: One of the most significant advantages of Diffence is its ability to protect models without requiring retraining. This is a game-changer for organizations with large, complex, and costly models already in production. Defenders can now deploy a robust MIA defense as a plug-and-play component, significantly enhancing the privacy posture of their existing AI assets with minimal disruption or capital expenditure. This eliminates the prohibitive barrier of retraining, which has historically limited the adoption of many privacy-preserving techniques.
- Enhanced Privacy with Preserved Utility: Diffence offers a superior privacy-utility trade-off compared to previous methods. By intelligently selecting reconstructed samples that maintain the original classification label, it ensures that the model's accuracy remains uncompromised. Simultaneously, the normalization of predicted logits through the precomputed range significantly reduces the information leakage that MIAs exploit. This means defenders no longer have to make severe compromises between model performance and data privacy. For instance, the reduction in attack AUROC from 79% to 56% (when combined with Selena) represents a substantial practical improvement in privacy.
- Layered Defense Strategy: Diffence is not a standalone solution that replaces existing defenses; rather, it complements them. Its pre-inference phase operation allows it to be seamlessly integrated with both training phase defenses (like DPSGD or Selena) and post-inference phase defenses (like MemGuard). This enables defenders to construct a multi-layered, robust defense architecture, combining the strengths of different techniques to achieve cumulative privacy benefits. This synergistic approach offers a more comprehensive and resilient security posture against MIAs.
- Resource Efficiency: By obviating the need for retraining, Diffence offers substantial savings in computational resources, time, and engineering effort. This makes advanced privacy protection accessible even to organizations with limited resources or tight deployment schedules. The optional nature of additional data further reduces logistical burdens.
- Robustness Against Adaptive Attackers: The evaluation explicitly considered an adaptive attacker who possesses full knowledge of the Diffence defense mechanism. The sustained effectiveness of Diffence under such strong adversarial assumptions provides a high degree of confidence in its real-world resilience. Defenders can be assured that deploying Diffence offers protection even against sophisticated adversaries who will attempt to bypass the defense.
- Adaptability to Defender Knowledge: The three scenarios for determining the logit range (Scenario 1: known members/non-members; Scenario 2: known members; Scenario 3: no knowledge) illustrate Diffence's adaptability. Even in the most challenging scenario where a defender has no prior membership knowledge, Diffence remains effective. This flexibility ensures that privacy enhancements can be implemented across a wider range of operational contexts, regardless of the availability of auxiliary data.
In essence, Diffence empowers defenders with a practical, effective, and flexible tool to significantly bolster the privacy of their machine learning models. It addresses critical limitations of prior work, enabling a more proactive and comprehensive approach to mitigating membership inference risks in deployed AI systems.
Key Takeaways
- Novel Pre-Inference Defense: Diffence introduces a new paradigm for Membership Inference Attack (MIA) defense by operating at the pre-inference phase, modifying inputs before they reach the target model.
- No Retraining, Plug-and-Play: It eliminates the need for costly model retraining, making it a practical, plug-and-play solution for protecting already deployed machine learning models.
- Diffusion Models for Input Reconstruction: Diffence leverages diffusion models to reconstruct input samples, ensuring the target model never processes exact replicas of its training data, thereby reducing membership leakage.
- Enhanced Privacy-Utility Trade-off: The defense significantly improves the privacy-utility balance, reducing attack AUROC from 79% (undefended) to a new state-of-the-art of 56% when combined with other defenses like Selena, while preserving model accuracy.
- Intelligent Sample Selection: Diffence generates multiple reconstructed samples and selects the best one based on matching the original predicted label and ensuring the predicted logit falls within a precomputed range, maximizing both utility and privacy.
- Seamless Integration: It can be seamlessly integrated with existing training-phase (e.g., DPSGD, Selena) and post-inference-phase (e.g., MemGuard) defenses, enabling a powerful, multi-layered security strategy.
About the Speaker(s)
Yuefeng Peng is the primary speaker for "Diffence: Fencing Membership Privacy With Diffusion Models," presenting this collaborative work. He is affiliated with Amherst, presumably UMass Amherst, where he conducted this research alongside co-authors Ali and his advisor, Amir (likely Amir Houmansadr). His work centers on addressing critical privacy challenges in machine learning, specifically focusing on developing robust defenses against Membership Inference Attacks using advanced generative models like diffusion models.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic ML privacy research with a clean core idea — use diffusion model reconstruction as a pre-inference membership inference defense. Solid experimental coverage and a genuinely useful plug-and-play framing, but the privacy gains (AUROC 79% → 56%) are incremental rather than transformative, and the talk reads as a conference paper presentation rather than a research event.
Heather Calloway (CISO) — PASS
Technically competent academic work on a real privacy problem, but this is a research paper read aloud at a security symposium — not a talk for defenders, executives, or governance decision-makers. The gap between the attack surface and any institutional action is never crossed.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025