Rectifying Privacy and Efficacy Measurements in Machine Unlearning: A New Inference Attack Perspective
Nima Naderloui
34th USENIX Security Symposium (USENIX Security '25) · Day 3 · ML and AI Privacy 2
Overview
This talk, presented by Nima Naderloui, addresses critical flaws in the current evaluation frameworks for machine unlearning algorithms. While the field has seen an "explosion" of inexact unlearning methods, often showing incremental improvements, Naderloui argues that this apparent success stems not from the methods themselves, but from limitations in how they are assessed. The research highlights that existing evaluation metrics, primarily relying on average-case Membership Inference Attacks (MIAs), are inherently weak and fail to capture the true privacy leakage and unlearning efficacy.

Key moments
- 0:00 Introduction: Unlearning evaluation is tricky, is success real?
- 2:00 Identifying three key limitations in current unlearning evaluation
- 4:00 The gold standard: Retraining from scratch; defining efficacy and privacy
- 5:50 Privacy challenge: Non-uniform memorization and high-risk samples
- 7:40 Efficacy challenge: Output suppression leads to weak signals
- 9:00 Visualizing why average MIAs fail with vulnerable samples
Rectifying Privacy and Efficacy Measurements in Machine Unlearning: A New Inference Attack Perspective
Speakers: Nima Naderloui
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=w851hfRWwbs
Overview
This talk, presented by Nima Naderloui, addresses critical flaws in the current evaluation frameworks for machine unlearning algorithms. While the field has seen an "explosion" of inexact unlearning methods, often showing incremental improvements, Naderloui argues that this apparent success stems not from the methods themselves, but from limitations in how they are assessed. The research highlights that existing evaluation metrics, primarily relying on average-case Membership Inference Attacks (MIAs), are inherently weak and fail to capture the true privacy leakage and unlearning efficacy.
The core motivation for machine unlearning stems from privacy regulations like GDPR and CCPA, which grant individuals the Right to be Forgotten. As retraining models from scratch is computationally prohibitive for every data removal request, efficient inexact unlearning algorithms have emerged as a practical alternative. However, without robust evaluation, service providers may be releasing "unlearned" models that still inadvertently leak sensitive information or fail to effectively remove data. This talk introduces a novel framework, Roly, designed to provide a more accurate and granular assessment of unlearning algorithms, focusing on high-risk samples and calibrated measurements.
The work is crucial for the future of privacy-preserving machine learning. By exposing the shortcomings of current evaluation paradigms, it forces a re-evaluation of what constitutes successful unlearning. The findings suggest that many widely-acclaimed unlearning algorithms might not be as effective or privacy-preserving as previously believed, especially when subjected to more sophisticated, targeted attacks. This research provides a roadmap for developing and validating truly privacy-compliant unlearning solutions, ensuring that the promises of the Right to be Forgotten are genuinely met.
Background
▶ Watch: Introduction: Unlearning evaluation is tricky, is success real? (0:00)
The concept of machine unlearning has gained significant traction in recent years, driven primarily by evolving data privacy regulations such as the General Data Protection Regulation (GDPR) in Europe and the California Consumer Privacy Act (CCPA) in the United States. These regulations enshrine the Right to be Forgotten, allowing individuals to request the deletion of their personal data from systems, including those used for training machine learning models. For large-scale models trained on vast datasets, completely retraining a model from scratch every time a user requests data removal is computationally infeasible and resource-intensive. This practical constraint has led to the development of inexact unlearning algorithms, which aim to approximate the state of a model retrained from scratch, but at a significantly lower computational cost.
Prior work in machine unlearning has largely focused on demonstrating incremental improvements in these inexact algorithms. Evaluations typically rely on Membership Inference Attacks (MIAs), where an adversary attempts to determine if a specific data point was part of the model's training set. For unlearning, MIAs are adapted to distinguish if a sample has been successfully "unlearned" or if it was never part of the training data. Metrics such as MIA accuracy or True Positive Rate (TPR) at low False Positive Rate (FPR) are commonly reported. For instance, some state-of-the-art inexact unlearning algorithms, like Scrub and Sparse Fine-tuning, have reported MIA accuracies below 10%, suggesting high privacy preservation.
However, Naderloui and his collaborators argue that these existing evaluation frameworks suffer from three key limitations. Firstly, most works consider average-case MIAs, such as population-level MIAs or those using general test data. These approaches are inherently weaker as they average performance across all samples, including those that are easily generalized and thus less prone to memorization. Secondly, many evaluations run MIAs on only a small, randomly selected fraction of "forget data." A more realistic and adversarial scenario would involve an attacker targeting highly memorized or high-risk samples, which are more likely to reveal sensitive information. Thirdly, comparisons with models retrained from scratch are often incomplete. Two models might exhibit very similar overall accuracy, but their behavior can differ significantly on specific, sensitive samples. The "golden standard" for unlearning, as emphasized in the talk, remains a model entirely retrained from scratch without the removed data. This serves as the reference point for both efficacy (how well the model performs without the data) and privacy (how little information about the removed data is leaked).
To address these issues, the researchers reframe the evaluation through the lens of MIAs by defining four key data distributions:
- In: Samples that were part of the initial training data.
- Out: Samples that were explicitly excluded from the training data (i.e., the model was trained without them).
- Unlearn: Samples that were initially part of the training data but were subsequently targeted for unlearning by an algorithm.
- Held Out: Samples that were never part of the training data and were not unlearned (i.e., a model trained from scratch without them).
Informally, privacy is preserved if the Unlearn distribution is close to the Held Out distribution, meaning the unlearned model behaves as if the data was never seen. Efficacy is achieved if the Unlearn distribution is close to the Out distribution, meaning the unlearned model effectively removes the influence of the data. The closer these distributions, the better the unlearning performance. The challenge lies in accurately measuring these distances, especially when current methods are demonstrably insufficient.
Key Findings
▶ Watch: The gold standard: Retraining from scratch; defining efficacy and privacy (4:00)
The research presented in this talk reveals several critical insights into the evaluation of machine unlearning, challenging the prevailing optimism in the field:
- Current Evaluations Overestimate Unlearning Success: The most significant finding is that existing, average-case Membership Inference Attacks (MIAs) used to evaluate unlearning algorithms are fundamentally flawed. They consistently overestimate the privacy preservation and efficacy of these algorithms. By averaging across a population of samples, these MIAs fail to detect subtle but significant leakages on individual, vulnerable data points.
- Memorization is Not Uniform; Vulnerable Samples are Key: The talk highlights that models do not memorize all training data equally. Some samples are far more "vulnerable" to being memorized, making them high-risk for privacy leakage even after unlearning. Evaluating unlearning performance using random samples gives undue credit for protecting data that was already well-generalized and safe. The study strongly advocates for focusing evaluations on these high-risk, vulnerable samples to accurately assess privacy.
- Output Suppression Hides Efficacy Failures: A crucial observation is the phenomenon of output suppression in unlearned models. When an unlearning algorithm attempts to remove a sample, it often suppresses the model's output signals (e.g., logits or confidence scores) for that specific sample. While this might appear to indicate successful unlearning, it merely makes the sample indistinguishable from a truly removed sample due to a general lack of strong signal. This suppression makes it difficult to measure true unlearning efficacy, as the model's behavior on the unlearned sample cannot be directly compared to a model that genuinely never saw the data.
- Canary Injection Reveals Significant Privacy Leakage: The researchers demonstrate the power of canary injection — intentionally inserting known vulnerable samples into the forget request. When these canaries are used as the target for unlearning, the Roly framework reveals a significantly higher privacy leakage compared to unlearning random samples. For instance, injecting canaries as 50% of the forget data (and even at lower percentages like 5%) showed a more pronounced privacy risk, underscoring the importance of adversarial testing scenarios.
- Roly Framework Uncovers High MIA Success Rates: The proposed Roly framework provides a more granular, per-sample, and calibrated evaluation. Using Roly, the researchers found a dramatic increase in MIA success rates. For example, on Tiny ImageNet, Roly revealed up to 69% MIA success in distinguishing between unlearned and retrained samples. This is a stark contrast to previous benchmarks that reported much lower MIA accuracies (e.g., less than 3% or even below 10%), indicating that unlearning is far from a solved problem.
- Privacy Auditing is Distinct from Unlearning Auditing: The talk clearly differentiates between privacy auditing (does this model leak information?) and unlearning auditing (is this model indistinguishable from a retrained model?). Conflating these two distinct goals can lead to misinterpretations of an algorithm's true performance. An algorithm might reduce privacy leakage but still be distinguishable from a retrained model, indicating a lack of full efficacy.
- Privacy Leakage Does Not Always Respond to Efficacy: An interesting observation from experiments with sparse fine-tuning was its success in mitigating MIA attacks by maintaining low efficacy. While sparse fine-tuning achieved a low TPR of 10% at 1% FPR (suggesting privacy), it also showed an 80% success rate in general, implying it struggled with efficacy. This demonstrates that an algorithm might appear privacy-preserving by simply suppressing outputs, even if it hasn't truly removed the data's influence, further emphasizing the need for calibrated efficacy measurements.
Technical Deep Dive
▶ Watch: Privacy challenge: Non-uniform memorization and high-risk samples (5:50)
The core of this research lies in identifying the technical shortcomings of existing unlearning evaluation methods and proposing a robust framework to overcome them. The speaker meticulously details how current Membership Inference Attacks (MIAs) fall short and introduces the Roly framework as a more accurate alternative.
Limitations of Current MIAs
Existing evaluations, as highlighted by Naderloui, typically suffer from three critical technical flaws:
- Average-Case MIAs: Many works employ population MIAs or run MIAs on general test datasets. These methods average the attack accuracy across a broad population of samples. However, the degree of memorization varies significantly across different data points. Some samples are easily generalized by the model and their presence in the training set has minimal impact on the model's behavior, making them "safe." Other samples, often unique or complex, are highly memorized and thus "high-risk." Average-case MIAs dilute the signal from these vulnerable samples, giving a false sense of security regarding unlearning effectiveness. The talk visually illustrates this: distributions of 500 random samples are hard to distinguish with a linear model, whereas 500 vulnerable samples show a much clearer separation between their "unlearn" and "held out" states.
- Randomly Selected Forget Data: Adversaries, in a real-world scenario, would not randomly select data to infer membership. Instead, they would target samples they suspect are highly memorized or particularly sensitive. Current evaluations often use small fractions (e.g., 1-10%) of randomly chosen data as forget requests. This approach fails to simulate a worst-case attack where an adversary focuses on retrieving information about the most vulnerable data points.
- Incomplete Retrained Model Comparisons: While the ideal "golden standard" for unlearning is a model retrained from scratch without the forgotten data, comparisons are often superficial. Models might have similar overall accuracy metrics, but their internal representations or specific predictions for individual samples could still differ significantly. This subtle divergence, especially on forgotten samples, is often missed by aggregated metrics.
The Roly Framework
To address these limitations, the researchers propose Roly, a novel framework for rectifying unlearning efficacy and privacy evaluations. Roly is built on three foundational principles:
- Per-Sample Evaluation: Instead of aggregated metrics, Roly focuses on evaluating the unlearning process for individual samples. This allows for the identification of specific vulnerable data points and a more accurate assessment of per-sample privacy leakage and efficacy.
- Focus on Vulnerable Samples: Roly specifically targets high-risk, vulnerable samples for evaluation. This is achieved through canary injection, where known vulnerable samples are deliberately included in the forget request. This approach, inspired by privacy auditing from a game-theoretic perspective, ensures that the evaluation measures protection for data that genuinely needs it. The speaker gives the example of "Jamie" as a vulnerable sample versus "Alex" who is well-generalized.
- Calibration for Output Suppression: A critical technical challenge is output suppression. Unlearned models often produce weak or suppressed signals (logits, confidence scores) for forgotten samples. This suppression can make it appear as if the data has been removed, when in reality, the model simply isn't confidently predicting anything for it. To counter this, Roly introduces a simple calibration test model for efficacy evaluation. This test model acts as a "router," presenting an adversary with queries that are 50% from the unlearned model and 50% from a truly retrained model (without the data). The adversary does not know the source. This setup, although not existing in practice, allows for a fair, calibrated comparison of signals, enabling trusted users or model providers to accurately assess efficacy without being misled by output suppression.
Technical Implementation of Roly
The Roly framework relies on shadow models to estimate per-sample distributions. The process involves:
- Shadow Model Training: Multiple shadow models (
n) are trained to generate a distribution of outputs (e.g., log-likelihoods or logits) for each target sample under different training conditions (In, Out, Unlearn, Held Out). - Sample Splitting: Target samples are split into three groups to ensure sufficient shadow models for all required distributions (e.g., if a sample is "In" for one shadow model, it's "Out" for another, and "Unlearn" for a third, etc.). This ensures that for any given sample, there are enough observations to characterize its behavior across the four distributions.
- Hypothesis Testing: Statistical hypothesis tests are then performed on these estimated per-sample distributions to determine if the "Unlearn" distribution is statistically indistinguishable from "Held Out" (for privacy) or "Out" (for efficacy).
Unlearning Algorithms Evaluated
The study applied Roly to two top-performing inexact unlearning methods:
- Scrub: This algorithm maximizes KL divergence between the original model's predictions and a target distribution for the forgotten data, effectively "pushing" the model's prediction away from the memorized information.
- Sparse Fine-tuning: This method aims to achieve unlearning by selectively fine-tuning a sparse subset of model parameters, often leading to a reduction in model efficacy while attempting to mitigate MIAs.
Experimental Datasets and Metrics
Roly was evaluated on both vision datasets (Tiny ImageNet, CIFAR-10) and text datasets (Wikex for language modeling). For vision, experiments involved fine-tuning small vision transformers and unlearning less than 1% of the training data. For text, the focus was on unlearning N-gram sequences using techniques like gradient ascent and negative preference optimization. The primary metrics remained MIA success rate, often reported as TPR at low FPR (e.g., 1% FPR), and direct comparisons of distribution gaps.
The technical contribution of Roly is its ability to move beyond simplistic aggregate metrics to a nuanced, per-sample, and calibrated evaluation, providing a far more accurate picture of an unlearning algorithm's true privacy and efficacy.
Demo / Proof of Concept
▶ Watch: Efficacy challenge: Output suppression leads to weak signals (7:40)
While the talk did not feature a live, interactive demonstration, it presented compelling experimental results that served as a robust proof of concept for the Roly framework. These results showcased Roly's ability to uncover significant privacy and efficacy shortcomings in existing machine unlearning algorithms that were previously masked by weaker evaluation methods.
Key experimental findings presented included:
- High MIA Success on Vision Models: The most striking result was on Tiny ImageNet, where Roly revealed an astonishing up to 69% MIA success rate in distinguishing between unlearned and retrained samples. This figure dramatically contrasts with previous benchmarks, which often reported MIA accuracies of less than 3% or even below 10% for state-of-the-art unlearning methods. This demonstration clearly illustrated that models deemed "unlearned" by conventional metrics still leak substantial information about the forgotten data when subjected to Roly's rigorous, per-sample evaluation. The speaker noted that this gap existed even when the unlearned and retrained models behaved very similarly on average accuracy.
- Impact of Canary Injection: The talk provided strong evidence for the effectiveness of canary injection in revealing privacy leakage. When vulnerable samples were deliberately injected as canaries into the forget data (e.g., comprising 50% of the forget set, and even at lower percentages like 5%), the privacy leakage measured by Roly was significantly more pronounced. This proved that focusing on high-risk samples, rather than random ones, is crucial for accurate privacy auditing. The visualization showed a clearer separation between the "unlearn" and "held out" distributions when vulnerable samples were targeted.
- Comparison with Average-Case MIAs: The presentation explicitly highlighted the noticeable gap between Roly's findings and those from representative implementations of average-case MIAs, such as population attacks using linear regression models. Roly consistently demonstrated higher sensitivity to unlearning failures, proving its superior capability in detecting subtle information leakage.
- Efficacy vs. Privacy with Sparse Fine-tuning: Experiments on CIFAR-10 with sparse fine-tuning provided an insightful case study. This method showed success in mitigating MIA attacks, achieving a low TPR of 10% at 1% FPR, which might suggest good privacy. However, the overall success rate was around 80%, indicating that while it might suppress signals (appearing private), it struggled with true efficacy (removing the data's influence). This finding underscores the distinction between privacy auditing and unlearning auditing and the necessity of calibrated efficacy measurements to avoid misleading conclusions.
- Generalization to Language Models: To demonstrate the generalizability of the framework, experiments were extended to the Wikex dataset for language modeling. The talk provided a real-world example of unlearning an N-gram sequence (specifically, the last seven tokens of a text referring to "foundation data of the program") using techniques like gradient ascent and negative preference optimization. This showed that the principles of Roly could be applied beyond vision tasks, although the speaker acknowledged the significant challenges and expenses associated with applying Roly to large language models.
In essence, the experimental results served as a powerful proof of concept, empirically validating the limitations of current unlearning evaluation methods and demonstrating Roly's superior capability in providing a more accurate and granular assessment of unlearning algorithms.
Defensive Implications
▶ Watch: Visualizing why average MIAs fail with vulnerable samples (9:00)
The findings from this research have profound implications for practitioners, researchers, and organizations involved in developing and deploying machine unlearning solutions, particularly in light of stringent privacy regulations like GDPR and CCPA.
- Do Not Trust Current Unlearning Benchmarks at Face Value: The most critical defensive implication is to approach existing benchmarks and reported successes of inexact unlearning algorithms with extreme skepticism. If an unlearning algorithm claims high privacy preservation based solely on average-case Membership Inference Attacks (MIAs) or evaluations on randomly selected forget data, it is likely overstating its capabilities. Defenders should demand more rigorous, per-sample evaluations.
- Prioritize Per-Sample Evaluation for High-Risk Data: Defenders must shift their focus from aggregate metrics to per-sample evaluation, especially for data that is prone to high memorization. Identify and characterize vulnerable samples within your datasets. These are the data points most likely to leak sensitive information even after unlearning. Any unlearning solution must demonstrate robust protection for these specific, high-risk samples.
- Implement Canary Injection for Auditing: To genuinely audit the privacy guarantees of an unlearning system, organizations should adopt canary injection. This involves deliberately introducing known vulnerable data points (canaries) into the training set and subsequently requesting their unlearning. The ability of an adversary (or an internal auditor using tools like Roly) to detect these unlearned canaries serves as a powerful, real-world measure of privacy leakage. If canaries are still detectable, the unlearning process is insufficient. This should be integrated into both the unlearning request pipeline and the overall training pipeline.
- Calibrate Efficacy Measurements: Be acutely aware of output suppression. An unlearned model might appear to have forgotten data simply because its output signals for that data are suppressed, not because the data's influence has truly been removed. Defenders need to employ calibrated evaluation techniques, similar to Roly's calibration test model, to distinguish between genuine data removal and mere signal suppression. Without proper calibration, efficacy will be overestimated, leading to models that still retain problematic biases or information.
- Distinguish Privacy Auditing from Unlearning Auditing: It is crucial for security teams and data privacy officers to understand that privacy auditing (assessing information leakage) is distinct from unlearning auditing (assessing indistinguishability from a retrained model). An algorithm might achieve one without the other. Both aspects are vital for compliance and robust security. A model might not leak information about a specific sample but still behave differently from a truly retrained model, indicating potential issues with its overall integrity or efficacy.
- Invest in Robust Evaluation Frameworks: Organizations relying on unlearning for compliance (e.g., Right to be Forgotten requests) should invest in or develop robust evaluation frameworks that incorporate the principles of Roly. This includes training shadow models to estimate per-sample distributions and employing sophisticated statistical tests. While expensive, this investment is necessary to ensure genuine compliance and avoid potential legal and reputational risks associated with data leakage.
- Address Challenges for Generative Models and LLMs: The talk acknowledges that applying Roly directly to generative models and Large Language Models (LLMs) is currently challenging due to their scale and the difficulty in defining a "retrained" model. For these complex architectures, defenders should explore alternative strategies such as estimating memorization and employing sequential training techniques as suggested in the paper, while research continues to adapt robust evaluation to these models.
In summary, defenders must move beyond superficial evaluations and adopt a more adversarial, granular, and calibrated approach to assessing machine unlearning. The era of accepting incremental improvements based on weak metrics is over; the focus must now be on verifiable and robust privacy and efficacy guarantees.
Key Takeaways
- Current evaluation frameworks for machine unlearning are fundamentally flawed, often relying on weak, average-case Membership Inference Attacks (MIAs) that overestimate privacy preservation and efficacy.
- Memorization is not uniform; evaluations must focus on vulnerable samples and employ canary injection to accurately assess privacy leakage in worst-case scenarios.
- The phenomenon of output suppression in unlearned models can mask true efficacy failures; calibration is essential to distinguish genuine data removal from mere signal suppression.
- The Roly framework offers a more robust, per-sample, and calibrated evaluation method, revealing significantly higher MIA success rates (e.g., up to 69% on Tiny ImageNet) than previously reported, indicating that unlearning is far from a solved problem.
- Privacy auditing (does the model leak information?) and unlearning auditing (is the model indistinguishable from a retrained model?) are distinct goals and should not be conflated.
- While applying Roly to generative models and Large Language Models (LLMs) is challenging and expensive, future work should explore methods like memorization estimation and sequential training to mitigate privacy risks in these complex systems.
About the Speaker(s)
Nima Naderloui is the presenter of this paper, "Rectifying Privacy and Efficacy Measurements in Machine Unlearning: A New Inference Attack Perspective." He is a key collaborator on this research, which was conducted in conjunction with colleagues from Yukan Illoch in events instead of technology and the Alibaba Group. His work focuses on the critical area of evaluating the effectiveness and privacy guarantees of machine unlearning algorithms.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid, technically rigorous work that exposes a genuine methodological failure in how the field validates machine unlearning — the core insight that average-case MIAs are structurally incapable of catching per-sample leakage on vulnerable data is correct, important, and not obvious. The 69% MIA success rate on Tiny ImageNet against algorithms previously benchmarked below 10% is the kind of result that should make people uncomfortable, which is exactly what good measurement research does.
Heather Calloway (CISO) — WEAK
Legitimate and technically credible research exposing real flaws in how machine unlearning is evaluated — but it stops well short of the institutional, governance, or compliance threshold that makes it actionable for security leaders. The regulatory hook is invoked but never landed.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)