Few-shot Unlearning
Youngsik Yoon, Jinhwan Nam, Hyojeong Yun, Jaeho Lee, Dongwoo Kim, Jungseul Ok
IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 4
Overview
In the rapidly evolving landscape of machine learning, the ability to selectively remove specific data's influence from a trained model—a process known as machine unlearning—has become increasingly critical. This talk, "Few-shot Unlearning," presented by Youngsik Yoon and collaborators from POSTECH at IEEE S&P, tackles a significant challenge within this domain: performing unlearning when only a limited number of target data samples are available. Traditional unlearning methods often assume complete access to either the original training dataset or the full set of data intended for erasure, which is often an impractical assumption in real-world scenarios due to memory constraints, privacy regulations (e.g., GDPR's "right to be forgotten"), or the sheer volume of data.

Key moments
- 0:00 Introduction to machine unlearning and its challenges
- 1:13 Defining few-shot unlearning and limitations of simple removal
- 2:40 Introducing unlearning 'intentions' for specific model behaviors
- 3:25 Overview of the three-step unlearning framework
- 5:14 Detailed strategies for different unlearning intentions (Standard, Privacy, Miscorrection)
- 7:44 Key results for class removal scenario demonstrating effectiveness
Few-shot Unlearning
Speakers: Youngsik Yoon, PhD Student; Jinhwan Nam; Hyojeong Yun; Jaeho Lee; Dongwoo Kim; Jungseul Ok
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=QBz6KfaLUx8
Overview
In the rapidly evolving landscape of machine learning, the ability to selectively remove specific data's influence from a trained model—a process known as machine unlearning—has become increasingly critical. This talk, "Few-shot Unlearning," presented by Youngsik Yoon and collaborators from POSTECH at IEEE S&P, tackles a significant challenge within this domain: performing unlearning when only a limited number of target data samples are available. Traditional unlearning methods often assume complete access to either the original training dataset or the full set of data intended for erasure, which is often an impractical assumption in real-world scenarios due to memory constraints, privacy regulations (e.g., GDPR's "right to be forgotten"), or the sheer volume of data.
The core contribution of this research lies in formalizing and addressing few-shot unlearning, where the desired model behavior on inaccessible samples is, by definition, underspecified. Beyond simply removing data, the authors introduce the crucial concept of "unlearning intentions," allowing for more nuanced and purposeful unlearning objectives such as correcting mislabeled data or enhancing privacy protection against membership inference attacks. This work proposes a novel three-step framework—comprising model inversion, filtration, and unlearning with intention—that effectively approximates the behavior of an oracle model (a model retrained from scratch without the target data) even with minimal target data access.
This research is highly significant for the security and privacy community. It offers practical solutions for deploying machine unlearning in resource-constrained or privacy-sensitive environments, making it a powerful tool for maintaining model integrity, compliance, and robustness. By enabling precise control over model behavior with limited information, few-shot unlearning paves the way for more adaptable and ethically sound AI systems.
Background
▶ Watch: Introduction to machine unlearning and its challenges (0:00)
The concept of machine unlearning emerged as a response to the practical and legal challenges associated with static, immutable machine learning models. As models are increasingly deployed in sensitive applications, the need to remove the influence of specific training data points becomes paramount. This could be due to data corruption, privacy requests (such as those mandated by GDPR or CCPA), or the need to mitigate biases introduced by specific subsets of data. The most straightforward, albeit computationally expensive, approach to unlearning is to retrain the model entirely from scratch without the data to be forgotten. This "oracle model" serves as the ideal benchmark for unlearning methods.
Existing unlearning methods typically aim to approximate this oracle model more efficiently. Many of these methods, however, operate under the assumption that either the complete original training dataset or a comprehensive list of all data points to be unlearned is fully accessible. This assumption often breaks down in real-world contexts. For instance, in online learning scenarios, data might be processed and then discarded to save storage space. In federated learning, data never leaves the client devices, making centralized access impossible. Furthermore, the sheer volume of data to be unlearned can be overwhelming; specifying every instance of a mislabeled category, for example, might be infeasible.
This talk specifically highlights the limitations of prior work when confronted with few-shot unlearning scenarios. Consider a classifier trained on a dataset where hundreds of images of a specific object (e.g., "cross 7") are consistently mislabeled as another category (e.g., "2"). If the unlearning request only provides a few examples of the mislabeled "cross 7," simply removing these few examples from the training data and retraining is insufficient. The retained dataset would still contain hundreds of other mislabeled "cross 7" instances, leading to the model still misclassifying them. Moreover, the intention behind the unlearning request is often more complex than mere deletion. A user might want to correct the mislabeling of "cross 7" to its true label "7," rather than just removing a few instances. Similarly, they might want to protect "cross 7" against privacy attacks, not just eliminate its presence. These nuanced intentions, coupled with limited data access, create a significant gap that existing unlearning techniques struggle to address. The challenge is not just what to unlearn, but how to unlearn it effectively and stably, especially when the desired post-unlearning behavior is ill-determined or underspecified by the few available target samples.
Key Findings
▶ Watch: Introducing unlearning 'intentions' for specific model behaviors (2:40)
The research presented in "Few-shot Unlearning" makes several significant contributions to the field of machine unlearning, addressing critical limitations of prior work:
- Formalization of Few-shot Unlearning with Explicit Intentions: The authors formally establish the problem of few-shot unlearning, acknowledging the practical constraints of limited access to target data. Crucially, they extend the concept of unlearning beyond simple data removal to encompass explicit "intentions." These intentions include:
- Standard Unlearning: Aiming to remove the influence of target-like data entirely.
- Privacy Protection: Shielding target data from privacy breaches, specifically membership inference attacks, by maintaining predictions but decreasing confidence or altering representations.
- Mislabel Correction: Rectifying mislabeled target data to their correct labels, even when the full extent of mislabeling is unknown. This is a novel and highly practical intention for improving model accuracy and fairness.
- A Novel Three-Step Framework: To achieve few-shot unlearning with diverse intentions, the talk introduces a robust, generalizable three-step framework:
- Model Inversion: A specialized technique designed to "invert" the trained model and generate synthetic data that resembles the target data domain, thereby addressing the issue of limited data accessibility.
- Filtration: A method to identify and filter out generated samples that are either of insufficient quality or not sufficiently similar to the few-shot target data, ensuring that only relevant synthetic data is used.
- Unlearning with Intention: A set of tailored strategies that leverage the augmented target dataset to impose the specific unlearning intention (standard, privacy, or mislabel correction) onto the model.
- Demonstrated Effectiveness Across Diverse Scenarios: The proposed method is rigorously evaluated across three distinct scenarios: class removal, privacy protection, and mislabel correction. The results consistently show that the framework can achieve near-oracle performance, even when provided with as little as 3% of the target data. This includes:
- Achieving near-zero accuracy for erased classes while maintaining high accuracy for retained classes, preventing catastrophic forgetting.
- Significantly reducing the success rate of membership inference attacks on target data without compromising overall model utility.
- Successfully correcting mislabeled data, highlighting that simple erasure is often insufficient for complex unlearning objectives.
These findings collectively demonstrate a significant advancement in machine unlearning, providing a practical and versatile solution for scenarios where data access is limited and unlearning objectives are multifaceted.
Technical Deep Dive
▶ Watch: Overview of the three-step unlearning framework (3:25)
The core of the "Few-shot Unlearning" solution is a sophisticated three-step framework designed to overcome the challenges of limited target data and diverse unlearning intentions.
Step 1: Model Inversion for Data Accessibility
The first and most critical step is model inversion. Since few-shot unlearning implies limited access to the target data, the framework must effectively generate synthetic data that approximates the characteristics of the data to be unlearned or corrected. This is crucial because, unlike traditional unlearning, there isn't a large pool of specific instances to remove or modify. The authors design a new inversion technique specifically tailored for unlearning, which leverages the existing trained model, the few available target samples, and a general prior on the data domain.
To achieve this, they propose two novel losses for the model inversion process:
- Target Loss: This loss measures the semantic discrepancy by reversing a few-shot target data. Its purpose is to ensure that the generated synthetic data strongly reflects the semantic features present in the limited target samples provided. This helps the inversion process focus on generating data relevant to the specific unlearning task.
- Argumentation Loss: This loss measures the robustness of the generated samples to various data augmentations. By encouraging robustness, the generated data becomes more diverse and representative of the target data distribution, preventing the model from overfitting to the few-shot samples and ensuring that the unlearning generalizes effectively.
The combination of these losses allows the system to generate high-quality, relevant synthetic data that effectively "fills the gap" left by the limited real target samples, making the subsequent unlearning steps feasible.
Step 2: Filtration for Identifying Target-Like Data
Once a pool of synthetic data is generated through model inversion, the filtration step refines this output. Not all generated samples will be of sufficient quality or relevance to the unlearning task. This step involves two main sub-processes:
- Quality Filtering: First, the framework filters out images where the quality generated by the inversion technique is insufficient. This is achieved by using the original trained model to identify and discard samples that yield low confidence predictions. Low confidence often indicates that the generated sample does not strongly resemble any known class or is poorly formed, making it unsuitable for training.
- Similarity Differentiation: Next, the system differentiates between generated samples that are truly "target-like" and those that are not. This is accomplished by measuring the distance in the embedding space between the given few-shot target data (the real, limited samples) and the generated samples. Samples whose embeddings are sufficiently close to those of the real target data are retained, ensuring that the augmented dataset is genuinely representative of the unlearning target.
This filtration step is vital for ensuring that the subsequent unlearning process operates on a clean, relevant, and high-quality set of augmented target data.
Step 3: Unlearning with Intention
The final step involves devising specific unlearning strategies that impose the learner's desired intention, utilizing the augmented target data set from the filtration step. The talk outlines three primary intentions:
- Standard Unlearning (Class Remover Scenario):
- Goal: To remove the influence of target-like data completely. For instance, if class 7 needs to be unlearned, the model should predict near-zero probability for class 7 for any input, while maintaining high accuracy for other classes.
- Strategy: The augmented target data is used in a way that encourages the model to forget or ignore the specific class. The evaluation focuses on achieving near-zero accuracy on the data to be erased and high accuracy on the data to be retained. This prevents catastrophic forgetting, where unlearning one class negatively impacts performance on others.
- Privacy Protection (Against Membership Inference Attacks):
- Goal: To protect target-like data against membership inference attacks, which determine if a data point was used in training based on prediction confidence. The aim is to decrease the confidence of predictions on target data without removing them entirely, thus making it harder for an attacker to infer membership.
- Strategy: Instead of removal, the target-like data are "re-labeled" or their output is adjusted. Two specific strategies are proposed:
ours defend: This approach combines the output of the original model with the output of an untrained model (a model initialized but not trained). Specifically, it uses a weighted sum of the original model's output and the untrained model's output. The idea is to "soften" the original model's confident predictions for target data by introducing the uncertainty of an untrained model.ours untrained: This simpler strategy directly replaces the output for target-like data with the output of an untrained model. This completely nullifies any learned patterns for these specific data points, making them indistinguishable from unseen data.- Evaluation: Performance is assessed by measuring the attack success rate of membership inference attacks. A successful unlearning method for privacy protection should exhibit a significantly lower attack success rate on target data while maintaining overall model accuracy.
- Mislabel Correction:
- Goal: To correct the labels of mislabeled target-like data. This is particularly useful when specific categories have been consistently misclassified in the original dataset (e.g., MNIST cross 7 mislabeled as 2, needing correction to 7).
- Strategy: Again, the target-like data are not removed but relabeled according to the unlearning intention. Two strategies are presented:
ours correct: Used when the true label of the mislabeled data is known. The augmented target data are simply relabeled with their correct label, and the model is adjusted to learn this correction.ours negative: Used when the true label is unknown, but the intention is to simply adjust the "row" (presumably the logit or probability) corresponding to the mislabeled class into a "negative row," effectively suppressing the incorrect prediction. This is useful for general "bad label" correction without needing the precise true label.- Evaluation: The effectiveness is measured by the accuracy on the unlearned model for both the corrected data and the retained data. The goal is to show a substantial improvement in accuracy for the previously mislabeled samples.
Through these meticulously designed steps, the few-shot unlearning framework offers a versatile and powerful approach to modifying model behavior with precision, even under severe data constraints.
Demo / Proof of Concept
▶ Watch: Detailed strategies for different unlearning intentions (Standard, Privacy, M... (5:14)
While the talk did not feature a live, interactive demonstration, the authors presented a comprehensive experimental evaluation across three distinct scenarios, serving as a robust proof of concept for their few-shot unlearning framework. These evaluations directly compared their proposed methods against an "oracle model" (retrained from scratch) and existing unlearning techniques, both with full and limited access to target data. Crucially, they reported performance when only 3% of the target data was provided, showcasing the few-shot capability.
Scenario 1: Class Remover (Standard Unlearning)
In the class remover scenario, the objective was to completely remove all data belonging to a specific class (e.g., class 7) from the model's knowledge base. The ideal behavior for the unlearned model would be to predict a near-zero probability for class 7 for any input, while accurately classifying samples in all other classes.
The results demonstrated that their method, with both standard and alternate intentions, consistently achieved near-zero accuracy for data to be erased (class 7) regardless of the quantity of initial target data provided (even with just 3%). Simultaneously, it maintained high accuracy for the data to be retained. This outcome is significant because it indicates that the proposed model inversion method successfully generated suitable images for preventing catastrophic forgetting, a common issue where unlearning one piece of information inadvertently degrades performance on unrelated tasks. The model effectively forgot the target class without impairing its performance on other learned classes.
Scenario 2: Privacy Protection
The privacy protection scenario aimed to make a specific class (e.g., cross 7) safe against membership inference attacks. The evaluation used the attack success rate of a membership inference attack as the primary metric. A successful unlearning method in this context would result in a low attack success rate for the data to be erased, while maintaining the model's overall accuracy.
The privacy-oriented unlearning methods, specifically ours untrained and ours defend, exhibited significantly lower success rates on data to be erased compared to ours standard (which performs simple removal). This demonstrates their effectiveness in making it harder for an attacker to infer whether a specific data point from class 7 was part of the training set. Critically, these methods achieved this reduction in attack success rate while simultaneously keeping accuracies high on both the erased and retained data, confirming that privacy was enhanced without a significant loss in model utility.
Scenario 3: Mislabel Correction
The mislabel correction scenario addressed a practical problem: correcting mislabeled data. The authors used a synthetic SPO MNIST dataset where images of "cross 7" were intentionally mislabeled as "2." The goal was to correct these mislabeled "cross 7" instances to their true label, "7." The ideal unlearning outcome would be high accuracy on both the data to be retained and the corrected data.
A substantial gap was observed between ours correct (which relabels with the true label) and ours standard (which simply erases). Ours correct achieved significantly higher accuracy on the previously mislabeled data, effectively correcting the model's understanding of "cross 7." This finding strongly implies that merely erasing the few-shot mislabeled data is insufficient to remedy the misbehavior induced by the hundreds of other mislabeled instances remaining in the dataset. True correction requires a more active, intention-driven approach, as provided by their framework.
Across all scenarios, the evaluations consistently highlighted the superiority of their few-shot unlearning framework, particularly its ability to handle complex intentions and operate effectively with extremely limited target data, often outperforming existing methods that require more extensive data access.
Defensive Implications
▶ Watch: Key results for class removal scenario demonstrating effectiveness (7:44)
The research on few-shot unlearning has profound implications for defenders in various domains, from cybersecurity and privacy compliance to model robustness and ethical AI development.
- Enhanced Privacy Compliance: The most immediate implication is for compliance with data privacy regulations like GDPR's "right to be forgotten" or CCPA. Organizations often face requests to erase data, but fully retraining large models is prohibitively expensive and time-consuming. Few-shot unlearning provides a practical, efficient mechanism to remove the influence of specific user data, even when only a few examples of that user's data are accessible, significantly easing the burden of compliance. The privacy protection intention specifically strengthens models against membership inference attacks, a growing concern for sensitive datasets.
- Mitigating Model Bias and Misinformation: The mislabel correction intention offers a powerful tool for addressing model biases or inaccuracies stemming from mislabeled training data. If a small number of critical examples reveal a systemic mislabeling issue (e.g., a particular demographic consistently misclassified), defenders can use few-shot unlearning to correct these errors without requiring a full re-annotation of the entire dataset. This can lead to more fair, accurate, and trustworthy AI systems.
- Robustness Against Data Poisoning: While not explicitly a focus of the talk, the ability to selectively unlearn with limited information could be leveraged to mitigate certain types of data poisoning attacks. If an attacker injects malicious data points that subtly influence model behavior, and a few examples of this poisoned data are identified, few-shot unlearning could be used to neutralize their impact without needing to rebuild the entire training set.
- Cost-Effective Model Maintenance: For large-scale models, retraining from scratch is computationally expensive. Few-shot unlearning offers a more agile and cost-effective approach to model updates. Whether it's removing outdated information, correcting newly discovered errors, or adapting to changing privacy requirements, this method allows for precise model modifications without incurring the full cost of a complete retraining cycle.
- Improved Model Interpretability and Control: By explicitly defining "intentions" for unlearning, defenders gain a higher degree of control and interpretability over how models are modified. This moves beyond simple data deletion to purposeful behavioral adjustments, enabling engineers to more finely tune model responses to specific inputs or categories.
- Addressing Data Accessibility Limitations: The framework's ability to operate with only 3% of target data is a game-changer for scenarios where data access is inherently limited, such as in federated learning, online learning, or environments with strict data governance policies. Defenders can still implement effective unlearning strategies even when the full original dataset is unavailable.
In summary, defenders should strongly consider adopting intention-driven few-shot unlearning as a critical component of their machine learning lifecycle management. It offers a versatile, efficient, and robust mechanism for maintaining model integrity, ensuring compliance, and enhancing overall system security and fairness in an increasingly complex AI landscape.
Key Takeaways
- Few-shot unlearning is crucial for real-world scenarios: Traditional machine unlearning often assumes full data access, which is impractical due to memory costs, privacy concerns, and overwhelming target data volumes. Few-shot unlearning addresses this by enabling effective unlearning with only a small fraction (e.g., 3%) of the target data.
- Unlearning requires explicit intentions: Beyond simple data removal, effective unlearning necessitates clearly defined goals such as standard data removal, privacy protection against membership inference attacks, or precise mislabel correction.
- A three-step framework enables effective few-shot unlearning: The proposed framework consisting of model inversion (to generate synthetic data), filtration (to select high-quality, target-like data), and unlearning with intention (to apply specific strategies) provides a robust solution.
- Model inversion prevents catastrophic forgetting: The specialized model inversion technique is key to generating suitable images for the unlearning process, ensuring that removing the influence of target data does not negatively impact the model's performance on retained data.
- The framework is versatile and robust: It successfully demonstrates effectiveness across diverse and practical scenarios including class removal, enhancing privacy against membership inference attacks, and correcting mislabeled data (e.g., MNIST cross 7 mislabeled as 2, corrected to 7).
- Simple erasure is often insufficient: For complex unlearning objectives like mislabel correction, merely erasing a few target samples is not enough; intention-driven strategies are necessary to remedy systemic misbehaviors induced by the data.
About the Speaker(s)
The talk "Few-shot Unlearning" was presented by Youngsik Yoon, a PhD student at POSTECH (Pohang University of Science and Technology). He collaborated on this joint work with Jinhwan Nam, Hyojeong Yun, Jaeho Lee, Dongwoo Kim, and Jungseul Ok. The presentation highlights the research contributions from this team, focusing on advanced machine learning techniques for data privacy and model control.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This work formalizes few-shot unlearning with critical explicit intentions (privacy, mislabel correction), addressing a major real-world challenge. The novel three-step framework, particularly the model inversion and intention-driven strategies, achieves near-oracle performance with minimal data, a significant practical advancement for compliance and model integrity.
Heather Calloway (CISO) — STRONG ACCEPT
This research offers a critical, practical framework for machine unlearning under real-world data constraints, directly addressing pressing governance and compliance challenges like the "right to be forgotten." Its emphasis on specific unlearning "intentions" provides security leaders with a powerful new tool for managing model risk and ensuring institutional accountability.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024