Attributing Open-Source Contributions is Critical but Difficult: A Systematic Analysis of GitHub Practices and Their Impact on Software Supply Chain Security

Jan-Ulrich Holtgrave

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Github + OSN Security

Overview

In an era of increasing data privacy regulations, the "right to be forgotten" has become a critical challenge for machine learning systems. This technical article synthesizes key insights from a session at the NDSS Symposium, which explored the burgeoning field of machine unlearning. The talks presented diverse yet complementary approaches to ensuring data privacy, intellectual property protection, and model hygiene by effectively removing the influence of specific data points or experiences from trained models.

Watch on YouTube · Slides

Key moments

  1. 8:30 Introducing unlearning in offline reinforcement learning and 'Trash Deleter'
  2. 9:15 Motivations for unlearning: privacy, sensitive data, GDPR compliance
  3. 10:00 Limitations of current unlearning solutions: retraining, fine-tuning, random reward
  4. 11:45 Proposed method: Trajectory Deletor for stable unlearning
  5. 13:15 Introducing convergence training for guaranteed agent stability
  6. 14:30 Auditing unlearning: Verifying data deletion using inference attacks
  7. 16:00 Experimental design, investigated algorithms, and evaluation metrics

Machine Unlearning: Protecting Data and Privacy in Reinforcement Learning and Beyond

Speakers: Ching Gong (University of Virginia), Dion G (University of Technology Sydney), Deiri Wang (CSRO's Data 61, University of Chicago)

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=YWXJ0gh70p4

Overview

In an era of increasing data privacy regulations, the "right to be forgotten" has become a critical challenge for machine learning systems. This technical article synthesizes key insights from a session at the NDSS Symposium, which explored the burgeoning field of machine unlearning. The talks presented diverse yet complementary approaches to ensuring data privacy, intellectual property protection, and model hygiene by effectively removing the influence of specific data points or experiences from trained models.

The session delved into the complexities of unlearning, particularly within Reinforcement Learning (RL) environments, where sequential decision-making and cumulative rewards introduce unique challenges. It also addressed the broader issue of data protection through provably unlearnable examples, offering a proactive defense against data exploitation. The discussions highlighted the urgent need for efficient, verifiable, and robust unlearning mechanisms that can meet regulatory demands and mitigate emerging threats to data ownership and model integrity.

The speakers, Ching Gong, Dion G, and Deiri Wang, presented their research on specific techniques for machine unlearning. Gong introduced a novel "Trajectory Deletator" for offline RL, G explored general reinforcement unlearning strategies, and Wang unveiled a framework for certified data unlearnability. Together, their contributions underscore the technical sophistication required to navigate the ethical and legal landscape of data usage in advanced AI systems.

Background

▶ Watch: Introducing unlearning in offline reinforcement learning and 'Trash Deleter' (8:30)

The concept of machine unlearning arises from a fundamental tension in modern AI: the desire to leverage vast datasets for powerful models versus the necessity to protect individual privacy and intellectual property. Machine unlearning aims to selectively remove the influence of specific training data from a model, making it behave as if that data had never been used in its training process. This capability is paramount for several reasons:

Firstly, legal frameworks such as the General Data Protection Regulation (GDPR) grant individuals the "right to erasure" or the "right to be forgotten." Compliance with such regulations requires mechanisms to permanently delete personal data and erase its impact on any derived models. Simply deleting the raw data is insufficient if its indelible mark remains within a deployed AI system.

Secondly, models can sometimes be trained on erroneous, outdated, or sensitive information that later needs to be expunged. This could involve patient sensitive information in healthcare applications, proprietary room layouts in robotics, or even incorrect historical data that could bias future decisions. The ability to remove such information without full retraining is crucial for maintaining model integrity and ethical operation.

Thirdly, the rise of powerful foundation models, particularly large language models (LLMs) and large vision models, has amplified concerns around intellectual property (IP) infringement. Artists, writers, and creators are increasingly finding their works used without explicit consent in model training, leading to models that can mimic their unique styles. Unlearning offers a potential recourse for protecting IP by removing the influence of specific creative works.

Traditional approaches to removing data influence, such as retraining the model from scratch on the remaining data, are often prohibitively time-consuming and computationally expensive, especially for large-scale models or when unlearning requests are frequent. Other methods like fine-tuning on the remaining data have proven unreliable, as they may not fully erase the target data's impact. The challenge lies in developing methods that are both efficient and provide a strong guarantee of true unlearning.

The domain of Reinforcement Learning (RL) presents unique complexities for unlearning. RL agents learn through sequential interactions with an environment, optimizing for cumulative rewards based on observed states and chosen actions. The knowledge acquired is deeply embedded in the agent's policy and value functions, making it challenging to precisely excise specific "memories" or "experiences" (trajectories) without disrupting the agent's overall learned behavior. The inherent instability of some RL training processes further complicates this.

Beyond reactive unlearning, proactive data protection measures have emerged, such as unlearnable examples (UEs). These involve perturbing data points before publication to render them unlearnable by models, thereby reducing their utility for malicious purposes like Membership Inference Attacks (MIAs) or training adversarial domain experts. However, existing UE methods face problems with generalization across diverse models and training strategies, and are vulnerable to recovery attacks, where attackers can restore model performance on unlearnable data with minimal effort.

Key Findings

▶ Watch: Limitations of current unlearning solutions: retraining, fine-tuning, random ... (10:00)

The NDSS session on machine unlearning brought forth several significant contributions across the spectrum of unlearning research:

  1. Efficient and Stable Offline RL Unlearning: Ching Gong introduced the Trajectory Deletator, a two-stage approach for offline Reinforcement Learning that combines forget training and convergence training. This method drastically reduces the computational cost compared to full retraining, achieving similar unlearning results in just 2.2% of the time, while ensuring model stability and preventing unintended forgetting of legitimate data. The research also proposed using Membership Inference Attacks (MIAs) as an auditor to verify the success of unlearning.
  1. General Reinforcement Unlearning Strategies: Dion G presented two distinct methods for Reinforcement Unlearning (RLU): a Decremental RL-based approach and a Poisoning-based approach. Both strategies demonstrated effectiveness in making an agent forget specific environments by either adjusting its policy or modifying the environment's dynamics, respectively. Crucially, these methods maintained the agent's performance in other, non-target environments, showcasing their precision and efficiency over traditional baselines. The privacy study on recommendation systems underlined their practical utility.
  1. Provable Guarantees for Unlearnable Data: Deiri Wang unveiled a novel Certified Data Learnability Framework designed to establish an upper bound on the utility of pirate models trained on protected data. This framework addresses the limitations of prior unlearnable examples by generating Provably Unlearnable Examples (PUEs). PUEs, crafted with the help of Random Weight Perturbation (RWP) in surrogate models, offer enhanced robustness against diverse adversarial models and are significantly more resilient to sophisticated recovery attacks that could otherwise restore model performance.

Technical Deep Dive

▶ Watch: Proposed method: Trajectory Deletor for stable unlearning (11:45)

The three presentations at the NDSS Symposium offered distinct yet interconnected technical approaches to the multifaceted problem of machine unlearning, each tackling specific challenges within the broader domain.

Trajectory Deletator for Offline Reinforcement Learning (Ching Gong)

Ching Gong's work focused on offline Reinforcement Learning (RL), where agents are trained on a fixed dataset of pre-collected trajectories. These trajectories consist of sequences of state, action, and reward pairs, representing the agent's interactions with an environment. The motivation for unlearning in this context is often driven by privacy concerns, such as the need to delete sensitive patient information or comply with GDPR.

Gong highlighted the limitations of existing unlearning solutions:

  • Retraining: The most straightforward method, considered the "golden line" for comparison, involves dropping the unlearning dataset and retraining the original agent from scratch. While effective, it is time-consuming and expensive, especially for frequent unlearning requests.
  • Fine-tuning: A popular method in other data unlearning fields, it involves dropping the unlearning dataset and fine-tuning the agent on the remaining data. However, this process is often unreliable for true unlearning.
  • Random Reward: Making rewards random for specific state-action pairs could confuse the agent into ignoring certain information. Yet, this approach can make training very unstable.

To address these challenges, the Trajectory Deletator was proposed. This method comprises two critical, sequential stages:

  1. Forget Training: This initial stage aims to minimize the influence of the target trajectories. The core idea is to apply a "minimizing" objective function, in contrast to the "maximizing" objective of standard RL. However, directly using this forgetting training process can be unstable and lacks convergence guarantees for the agent.
  2. Convergence Training: This is the crucial second stage designed to stabilize the unlearning process. In this phase, the unlearned agent is trained to learn the value functions of the original agent, but critically, only on the remaining data set (i.e., data that was not marked for unlearning). This ensures that while the target trajectories are forgotten, the agent's overall performance and stability on legitimate data are preserved. A key theoretical finding supports that this process can make the agent converge to a unique optimal function or agent, provided the two stages are performed separately and not simultaneously.

A significant aspect of the Trajectory Deletator is its auditing mechanism to verify successful unlearning. Gong proposed using Membership Inference Attacks (MIAs) as an auditor. This involves comparing the features of shadow agents (known to include the unlearning trajectories) with the features of the unlearning agent. The inherent instability of offline RL training, which is typically a drawback, can actually be leveraged to improve the efficiency of the auditor. By fine-tuning original agents on the original dataset, multiple shadow models can be generated that are sufficiently different for effective comparison, a technique that is less effective in stable supervised learning contexts.

Experimental results demonstrated that the Trajectory Deletator achieved results comparable to full retraining in only 2.2% of the time. The importance of convergence training was highlighted, showing that without it, the agent would suffer a significant decrease in performance and even "forget" some legitimate trajectories from the remaining dataset. Hyperparameter tuning revealed that while the forgetting steps (e.g., 8,000) significantly impact unlearning performance, the convergence steps (e.g., 2,000) are less sensitive to the agent's overall performance, with a small number often being sufficient. This total of 10,000 steps represented only 1% of the typical retraining steps, underscoring the efficiency gains.

Reinforcement Unlearning Strategies (Dion G)

Dion G expanded on the general concept of Reinforcement Unlearning (RLU), defining it as the process of enabling an RL agent to forget specific training environments, including all previously learned behaviors, experiences, and knowledge, such that it behaves as if it encountered that environment for the first time. The motivations echoed similar privacy (GDPR), data hygiene (incorrect/outdated information), and adaptability concerns.

The problem was formally defined as updating a learned policy pi to a new policy pi' such that the accumulated rewards in an unlearned environment (mu) are minimized, while the rewards received in other environments remain unchanged. This emphasizes the need for targeted unlearning without collateral damage to other learned behaviors.

Dion G proposed two primary methods for RLU:

  1. Decremental RL-based Approach: This method directly adjusts the agent's policy. It involves two steps:
  • Exploration: The agent thoroughly explores the unlearning environment mu to collect a sufficient number of experience samples.
  • Fine-tuning with Loss Function: The agent is then fine-tuned in the unlearning environment using a specially designed loss function. This function has two main terms: the first guides the new policy pi' to operate efficiently (in a "forgotten" manner) within mu, while the second encourages pi' to maintain its performance in the remaining environments. This approach aims to gradually erase the agent's learned knowledge in mu.
  1. Poisoning-based Approach: This method directly modifies the learning environment's dynamics. It involves three iterative steps:
  • Random Poisoning: The transition function (T_u) of the unlearning environment is initially modified using a random poisoning strategy. This means that for a given state s and action a, the agent will observe a new, poisoned state s_hat_prime instead of the original s_prime.
  • Agent Learning: The agent then learns a new policy in this modified (poisoned) environment.
  • Strategy Update: Based on the agent's newly learned policy, the poisoning strategy is updated using a proposed reward function. This reward function quantifies the difference between the learned policy and the existing policy, and combines rewards received in the remaining environments. The iterative nature encourages the agent to learn new but incorrect knowledge regarding the unlearning environment, thereby effectively "forgetting" its original, correct understanding.

Comparing the two approaches, the decremental method is generally less computationally intensive as it relies on fine-tuning the agent's policy without altering the environmental dynamics. In contrast, the poisoning-based approach is more computationally intensive due to its iterative process of adjusting the poisoning strategy and modifying the environment itself. However, the poisoning approach is often more effective because it directly targets the environment's underlying structure, ensuring a more thorough unlearning process.

Experimental evaluations on platforms like Grid World, Virtual Home, and Maze Explorer demonstrated that both proposed methods successfully degraded the agent's ability in the unlearning environment, evidenced by needing more steps to complete tasks and receiving fewer rewards. Crucially, the agents' learned behavior in other environments remained largely unaffected. A privacy study using recommendation systems showed that for an unlearned user, recommendation accuracy plummeted from 92% to approximately 68% (decremental) and 64% (poisoning), while accuracy for remaining users stayed almost unchanged, confirming the methods' efficacy and precision. Both methods also proved to be much more computationally efficient than baseline retraining approaches.

Provably Unlearnable Data Examples (Deiri Wang)

Deiri Wang's presentation tackled the proactive protection of data using unlearnable examples. The motivation stemmed from the increasing risks posed by the learnability of publicly available data, including Membership Inference Attacks (MIAs), the training of substitute models for transferable adversarial attacks, and the fine-tuning of large foundation models (LLMs, vision models) to extract sensitive knowledge or infringe on intellectual property (e.g., an artist's style from just 20 paintings). A significant problem with existing unlearnable examples (UEs) is their vulnerability to recovery attacks, where an adversary can slightly perturb the weights of a pirate model trained on UEs and, with a small amount of clean data and projected SGD, restore the model's performance to near-normal levels.

To address these vulnerabilities, Wang introduced a Certified Data Learnability Framework designed to establish a provable upper bound on the utility of pirate models trained on protected data. The framework consists of three parts:

  1. Perturbation Generation: Given a dataset DS to be protected, the framework searches for a set of perturbations (delta). This is achieved with the help of a surrogate model, which itself is trained with Random Weight Perturbation (RWP). RWP is incorporated to make the surrogate model more suitable for later certification and to generate more robust Provably Unlearnable Examples (PUEs). The goal is to apply delta to DS such that it minimizes the training loss of the surrogate model.
  1. Certification: This is the core of the provable guarantee. The previously trained surrogate model's weights are randomized multiple times (N times) using Gaussian noise. For each randomized copy, its utility is evaluated using a chosen metric A and a test dataset. All these results are then used to compute QA learnability with the aid of a quantile parametric smoothing (QBS) function. The QBS function takes the randomized model parameters as input and computes the model utility at a given quantile Q among all attainable model utilities. The framework provides a guarantee on the maximum attainable utility (QA learnability) for any pirate model whose parameters fall within a certified parameter set (EA). This is a strict, mathematical upper bound on how much utility an adversary can extract.
  1. Generalization Learnability Score: For scenarios where the defender lacks access to a private test dataset, the framework allows sampling a test set from a closed domain to compute a generalization learnability score using Hinge-Span.

Key findings from this framework demonstrated that:

  • Using larger values of Q and EA in the certification process allows for certifying higher learnability scores.
  • RWP surrogate models are superior, producing higher certified learnability on the same set of unlearnable examples compared to baselines, indicating they create better certification surrogates.
  • PUEs (generated based on RWP surrogates) exhibit lower certified learnability compared to baselines under the same training method. This signifies that PUEs offer more robust protection against certifiable adversaries.
  • Beyond the certified parameter set, PUEs also provide stronger protection against general adversaries. Recovery attacks were found to be significantly harder against models trained on PUEs compared to other baselines (EMN and OPS), solidifying their robustness.

Demo / Proof of Concept

▶ Watch: Auditing unlearning: Verifying data deletion using inference attacks (14:30)

The efficacy of the proposed unlearning methods was empirically demonstrated across all three presentations, providing concrete evidence of their performance and practical utility.

Ching Gong's Trajectory Deletator for offline Reinforcement Learning showcased its efficiency by achieving comparable unlearning results to the "golden line" of full retraining, but in a mere 2.2% of the time. The experiments utilized standard metrics: position and recall to measure the accuracy of the unlearning rate, and average returns to quantify the agent's performance. The results clearly indicated that the two-stage approach (forget training followed by convergence training) was critical for both effective unlearning and maintaining the agent's overall utility on the remaining, legitimate data. Without the convergence training, the agent's performance saw a significant decrease, and unintended forgetting of non-target data occurred.

Dion G's Reinforcement Unlearning methods were rigorously evaluated across diverse environments, including Grid World, Virtual Home, and Maze Explorer. The demonstrations consistently showed that both the Decremental RL-based and Poisoning-based approaches successfully degraded the agent's performance in the designated unlearning environments. This was evidenced by the agent requiring more steps to complete tasks and receiving fewer cumulative rewards. Crucially, these performance degradations were localized to the unlearned environments, with the agent's behavior and reward accumulation remaining stable and effective in other, non-target environments. Furthermore, a privacy study using recommendation systems provided a compelling real-world example: for an unlearned user, the recommendation policy accuracy dropped dramatically from 92% to approximately 68% and 64% for the two methods respectively, while the accuracy for other users remained virtually unchanged. This highlights the methods' precision in targeted unlearning. The computational overhead analysis also demonstrated that both proposed methods were significantly more efficient than baseline retraining methods.

Deiri Wang's framework for Provably Unlearnable Data Examples (PUEs) provided experimental validation of its robustness and certification capabilities. The demonstrations showed that using Random Weight Perturbation (RWP) in the surrogate model for perturbation generation led to higher certified learnability scores, indicating a more effective mechanism for creating unlearnable data. More importantly, the PUEs themselves, when compared to other baselines (EMN and OPS), consistently resulted in lower certified learnability. This empirical finding directly supports the claim that PUEs offer more robust protection against adversaries. The framework also specifically evaluated protection against recovery attacks, demonstrating that models trained on PUEs were significantly harder to recover than those trained on traditional unlearnable examples. This crucial proof of concept highlights PUEs' resilience against sophisticated adversarial attempts to circumvent unlearning.

Defensive Implications

▶ Watch: Experimental design, investigated algorithms, and evaluation metrics (16:00)

The research presented in this session offers several critical implications for defenders, ranging from data owners and model developers to organizations navigating regulatory compliance:

  1. Prioritize Machine Unlearning for Regulatory Compliance: For any organization handling sensitive user data, particularly those subject to GDPR or similar "right to be forgotten" mandates, implementing robust machine unlearning capabilities is no longer optional. The Trajectory Deletator and general Reinforcement Unlearning approaches provide blueprints for efficiently and verifiably removing data influence from complex RL systems, which are increasingly prevalent in areas like personalized recommendations, autonomous systems, and healthcare.
  1. Proactive Data Protection with Provably Unlearnable Examples (PUEs): Data publishers, creators, and intellectual property owners should consider adopting the Certified Data Learnability Framework to generate Provably Unlearnable Examples (PUEs) before publishing data to public platforms. This proactive defense can significantly reduce the utility of their data for malicious purposes, such as unauthorized model fine-tuning, membership inference attacks, or style mimicry by generative AI, providing a strong data availability guarantee.
  1. Beware of Recovery Attacks on Traditional Unlearnable Examples: Model owners and those implementing unlearnable data examples must be aware of the potent Recovery Attack vulnerability. Traditional UEs may offer a false sense of security. Defenders should prioritize the use of PUEs which have been shown to be robust against such attacks, ensuring that data designated as unlearnable truly remains so, even against determined adversaries.
  1. Embrace Auditable Unlearning Processes: The concept of using Membership Inference Attacks (MIAs) as an auditor for unlearning, as proposed by Gong, is a crucial step towards building trust and verifiability in unlearning systems. Defenders should integrate similar auditing mechanisms to objectively confirm that unlearning requests have been successfully executed and that the model no longer retains significant influence from the target data.
  1. Leverage Multi-Stage and Environment-Modifying Unlearning: For RL systems, the two-stage approach of forget training and convergence training (Trajectory Deletator) is vital. The convergence training ensures stability and prevents unintended forgetting of legitimate data. Furthermore, Dion G's work demonstrates that unlearning can be achieved through diverse strategies—either by directly adjusting the agent's policy (Decremental RL-based) or by strategically modifying the environment's dynamics (Poisoning-based). Defenders should evaluate which approach is most suitable given their specific RL architecture and environment.
  1. Insights for Large Language Models (LLMs): The speakers noted that techniques like Reinforcement Learning from Human Feedback (RLHF) used in LLMs can be viewed as a specific type of offline RL. This suggests that the unlearning strategies developed for RL could provide valuable insights and foundational methods for addressing unlearning challenges in the rapidly evolving domain of large language models, offering avenues for better control over their training data and outputs.

Key Takeaways

  • Machine unlearning is an essential capability for modern AI systems, driven by privacy regulations (e.g., GDPR), the need for data hygiene, and intellectual property protection.
  • Efficient and stable unlearning in Reinforcement Learning (RL) is achievable through multi-stage approaches like the Trajectory Deletator, which combines targeted forgetting with crucial convergence training to maintain overall model utility.
  • Diverse strategies exist for Reinforcement Unlearning (RLU), including directly adjusting an agent's policy (Decremental RL-based) or perturbing the environment's dynamics (Poisoning-based), both proving effective and computationally efficient.
  • Provably Unlearnable Examples (PUEs) offer a robust, proactive defense against data exploitation, providing certified guarantees on data unlearnability and strong resilience against sophisticated recovery attacks.
  • Auditing mechanisms, such as Membership Inference Attacks (MIAs), are critical for verifying the success and completeness of unlearning processes, building trust and ensuring compliance.
  • The principles and techniques developed for unlearning in RL have significant implications for future research and implementation in large models like LLMs, particularly those leveraging RL-based training methods.

About the Speaker(s)

The session on machine unlearning featured three distinguished researchers presenting their work:

Ching Gong is affiliated with the University of Virginia. Gong presented the research on the "Trajectory Deletator," a novel method for efficient and stable unlearning in offline reinforcement learning, including an auditing mechanism based on membership inference attacks.

Dion G is from the University of Technology Sydney. G discussed the broader concept of "Reinforcement Unlearning" and introduced two distinct approaches: a decremental RL-based method and a poisoning-based method, demonstrating their effectiveness and efficiency in making RL agents forget specific environments while preserving other learned behaviors.

Deiri Wang is associated with CSRO's Data 61 and the University of Chicago. Wang presented the paper on "Provably Unlearnable Data Examples," detailing a framework for certified data learnability that introduces Provably Unlearnable Examples (PUEs) to provide robust protection against data exploitation and recovery attacks. The project was funded and supported by the cyber security cooperative research center of Australia.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Three competent academic contributions to the machine unlearning problem — RL-specific unlearning is legitimately underexplored, and the certified unlearnability framework is the most interesting piece here. None of it is groundbreaking enough to define the conversation, but it's real work done by people who clearly understand the problem space.

Heather Calloway (CISO) — WEAK

Technically credible academic work on machine unlearning across RL contexts, but the session never crosses the bridge from research to institutional responsibility. The gap between 'this algorithm achieves unlearning in 2.2% of retraining time' and 'here is how your organization handles a GDPR erasure request against a deployed recommendation model' is never closed.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025