DLBox: New Model Training Framework for Protecting Training Data

Jaewon Hur (Postdoc · Georgia Tech)

Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · ML Security

Overview

The proliferation of artificial intelligence, particularly deep learning, has led to an increasing demand for vast datasets to train sophisticated models. However, a significant hurdle in this ecosystem is the secure sharing of sensitive training data between data owners and AI developers. This talk introduces DLBox, a novel model training framework designed to protect proprietary and sensitive training data from unauthorized leakage during the model development process. Presented by Jaewon Hur from Georgia Tech, DLBox addresses the critical challenge of ensuring data utility for model training while simultaneously enforcing strict controls to prevent its misuse or exfiltration.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction and the data leakage problem
  2. 2:00 Various methods of AI data leakage
  3. 3:00 Introducing DLBox: A new training framework
  4. 4:15 Clarifying DLBox's goal: Preventing malicious leakage
  5. 5:00 Defining benign model training: The DGM rules
  6. 6:40 DLBox architecture: Sandbox and D Oracle components
  7. 7:40 DLBox workflow illustrated with a PyTorch example

DLBox: New Model Training Framework for Protecting Training Data

Speakers: Jaewon Hur (Postdoc, Georgia Tech)

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=_b4GlVgIJIc

Overview

The proliferation of artificial intelligence, particularly deep learning, has led to an increasing demand for vast datasets to train sophisticated models. However, a significant hurdle in this ecosystem is the secure sharing of sensitive training data between data owners and AI developers. This talk introduces DLBox, a novel model training framework designed to protect proprietary and sensitive training data from unauthorized leakage during the model development process. Presented by Jaewon Hur from Georgia Tech, DLBox addresses the critical challenge of ensuring data utility for model training while simultaneously enforcing strict controls to prevent its misuse or exfiltration.

The core problem DLBox tackles is the lack of systematic solutions to protect shared training data once it's handed over to AI developers. Current practices offer little to no guarantees that data will only be used for benign model training and not illicitly copied, embedded, or reverse-engineered. This vulnerability stifles data sharing, particularly in sensitive domains like healthcare where medical data is invaluable for AI research but carries significant privacy risks.

DLBox aims to foster data sharing by offering a trusted environment where data owners can lend their datasets without fear of unauthorized access or leakage. It achieves this by defining and enforcing a set of rules for "benign model training" within a secure framework, thereby allowing AI developers to train models while preventing malicious activities. This research is crucial for building trust in data-driven AI development and unlocking the full potential of collaborative machine learning.

Background

▶ Watch: Introduction and the data leakage problem (0:00)

The landscape of modern AI development frequently involves a symbiotic relationship between data owners, who possess valuable datasets, and AI developers, who have the expertise to build models but lack sufficient data. Consider a hospital (data owner) with extensive medical imaging data and an AI developer aiming to create a cancer detection model from X-ray images. The hospital wants to facilitate AI research but only permits data usage for model training, strictly prohibiting any other form of data access or exfiltration. In the current paradigm, once the data is transferred to the developer, the data owner effectively loses control. This fundamental trust gap is a major impediment to innovation and collaboration in deep learning.

The methods an AI developer, or a malicious actor masquerading as one, can employ to leak data are diverse and sophisticated. The simplest approach involves direct exfiltration: copying the dataset to an external hard drive or transmitting it over a network. Even in highly controlled, air-gapped environments, data can still be leaked through the trained model itself. Attackers can intentionally embed the training data directly into the model's parameters, later decoding it to reconstruct the original samples. Alternatively, they can bias the training procedure to make the AI model disproportionately resemble specific data points, allowing for reconstruction attacks. More advanced techniques, such as gradient inversion attacks, can reconstruct input data from model gradients. The speaker emphasizes that "they can do whatever you can imagine," highlighting the broad spectrum of potential threats.

The fundamental dilemma lies in the conflicting goals of model training and data protection. For a model to achieve reasonable performance on a given task, it inherently must "remember" some information about the training data. The very act of obtaining a performant model implies that some data-derived information has been transferred. This makes complete prevention of data leakage seemingly impossible. DLBox acknowledges this inherent tension, positing that the goal is not to eliminate all information flow from data to model – which is necessary for training – but rather to distinguish between legitimate, benign information transfer (learning patterns) and malicious, unauthorized data exfiltration (reconstructing specific samples). This distinction forms the bedrock of DLBox's approach to defining and enforcing "benign model training."

Key Findings

▶ Watch: Introducing DLBox: A new training framework (3:00)

Given the inherent conflict between training a performant model and completely preventing data leakage, DLBox's core innovation lies in defining what constitutes benign model training. The research posits that benign model training is fundamentally a statistical process of learning common patterns from a dataset, rather than memorizing individual samples. This crucial distinction leads to several key observations:

  1. Equal Importance: In a benign training process, each sample in the dataset should have equal importance to the model. No single sample should be prioritized or discriminated against over others.
  2. Fair Contribution: Consequently, all samples in the dataset should contribute fairly and equally to the model's learning process.

Based on these observations, DLBox introduces the DGM rules, a set of three specific criteria designed to determine whether a model training process is benign or malicious. These rules capture the essence of standard, non-malicious deep learning training procedures:

  • D (Data Augmentation): Data samples should go through the same augmentation process. This means that transformations applied to the input data (e.g., flipping images, adding noise) should be uniformly applied across all samples or batches, preventing an attacker from manipulating augmentation to embed specific data.
  • G (Gradient Computation): Gradients should be backpropagated from equally computed loss values. This implies that the loss function, which quantifies the model's error, should treat all samples or batches equally, preventing an attacker from crafting specific loss values for particular samples to bias the model.
  • M (Model Update): The model should be updated only by adjusting all gradients. Standard training aggregates gradients (e.g., by averaging) across a batch or the entire dataset before updating model parameters. This rule prevents an attacker from making targeted updates based on individual samples, which could facilitate data embedding or reconstruction.

These DGM rules are designed to align with typical model training procedures, where samples are processed uniformly, losses are computed consistently, and model updates are derived from aggregated information. By enforcing these rules, DLBox aims to ensure that the model learns general patterns rather than memorizing or being biased by specific training examples, thereby mitigating the risk of data leakage through the model itself.

Technical Deep Dive

▶ Watch: Clarifying DLBox's goal: Preventing malicious leakage (4:15)

DLBox is engineered as a two-component framework: the D-Sandbox and the D-Oracle. This architecture leverages confidential computing to create a secure environment for training, meticulously monitoring and controlling all operations on the sensitive data. The implementation discussed by the speaker utilizes PyTorch for real-world machine learning model support and AMD SEV-SNP (Secure Encrypted Virtualization-Secure Nested Paging) to protect the integrity and confidentiality of the D-Sandbox.

The D-Sandbox acts as an isolated, secure execution environment for the untrusted AI developer's code. Its primary function is to prevent illegal leakage of training data. This is achieved through a meticulous process of address space isolation and controlled data interaction:

  1. Address Space Isolation: In a typical PyTorch execution, all variables, including the sensitive training dataset and user-defined variables, reside in the same address space. This grants the AI developer full control, enabling direct data access and exfiltration. DLBox addresses this by segmenting the address space into two distinct regions:
  • Unsafe Address Space: Where the untrusted AI developer's code is launched. Direct access to the training data is prohibited here.
  • Safe Address Space: Where the training data is securely loaded and managed. This space is protected by the confidential computing enclave (AMD SEV-SNP).
  1. Proxy Variables: To enable the AI developer's code to interact with the data in the safe address space without direct access, DLBox introduces proxy variables. These proxies act as references to the actual data objects, allowing the developer to invoke operations on the data indirectly.
  2. Sanitized API: All operations on the data invoked via proxy variables must go through a sanitized API. This API intercepts and validates every data flow operation, ensuring that it adheres to predefined security policies and does not facilitate unauthorized data access or modification. When an operation is invoked, the sanitized API performs the computation on the actual data in the safe address space, and a new proxy variable is assigned to the resulting object, maintaining the isolation.

The D-Oracle operates in conjunction with the D-Sandbox, serving as the real-time monitor and enforcer of the DGM rules. While the D-Sandbox ensures that data cannot be directly accessed or exfiltrated, the D-Oracle scrutinizes the computational flow to verify that the model training process itself is benign:

  1. Operation Monitoring: The D-Oracle continuously monitors all operations performed on the data within the D-Sandbox. This includes data loading, augmentation, loss computation, gradient calculation, and model updates.
  2. DGM Rule Enforcement: As operations occur, the D-Oracle checks them against the DGM rules:
  • It verifies that data augmentation procedures are uniformly applied across all samples or batches.
  • It confirms that loss values are computed equally for all samples, preventing biased loss functions.
  • It ensures that model updates are derived from aggregated gradients (e.g., averaged across a batch), preventing individual sample-based updates.
  1. Conditional Model Release: Only when the D-Oracle confirms that all DGM rules have been consistently satisfied throughout the training process will the trained model be released to the AI developer. If any rule is violated, indicating a potentially malicious or non-benign training attempt, the model is not returned, thus preventing the leakage of data through a maliciously trained model.

By integrating confidential computing with a fine-grained monitoring and enforcement mechanism, DLBox creates a robust framework that not only prevents direct data exfiltration but also ensures the integrity of the training process itself, allowing for secure and controlled data sharing.

Demo / Proof of Concept

▶ Watch: DLBox architecture: Sandbox and D Oracle components (6:40)

The evaluation of DLBox focused on quantifying its security enhancement and performance overhead, particularly concerning the amount of data that could be leaked through a "seemingly trained model." The speaker clarified that "seemingly trained" refers to a model returned by the untrusted AI developer, who could have performed any operation, not necessarily benign training.

For security evaluation, the researchers measured the number of samples leaked by one transfer of such a model. They utilized three different image datasets, though specific dataset names were not detailed in the transcript. The evaluation compared three scenarios:

  1. Baseline (No Security Measure): In this scenario, with no protection mechanisms in place, an attacker (AI developer) could simply leak the data itself. The results showed that 100% of the samples could be leaked with just one transfer, as the developer had full, unrestricted access to the raw data. The speaker also mentioned that attackers could embed data into the model or bias the training to make the model resemble the data, later decoding or reconstructing it.
  2. ADO Protection (Adversarial Defense Optimization): Even with some existing adversarial defense optimization techniques, which might offer mild protection against certain types of reconstruction attacks, attackers could still leak a significant portion of the samples. While not 100%, this implies a substantial vulnerability remains.
  3. DLBox: With DLBox enforcing the DGM rules within its secure framework, the evaluation demonstrated a drastic reduction in data leakage. The results indicated that almost no samples were leaked through the model training process. This validates DLBox's effectiveness in preventing both direct data exfiltration (via D-Sandbox) and malicious data embedding/biasing within the model (via D-Oracle's DGM rule enforcement). The speaker highlighted that the criterion for similarity was based on visual inspection, showing reconstructed images that were "very similar" to the originals in baseline attacks, but this was prevented by DLBox.

Beyond security, the performance overhead of DLBox was also evaluated. The framework was tested by training both image and language models. The results showed that DLBox imposes an average overhead of approximately 4% on training time. This indicates that the robust security guarantees provided by DLBox come with a relatively low computational cost, making it practical for real-world deployment in many machine learning scenarios. This balance between strong security and acceptable performance is critical for the adoption of such a framework.

Defensive Implications

▶ Watch: DLBox workflow illustrated with a PyTorch example (7:40)

DLBox represents a significant step forward in securing the deep learning supply chain, particularly concerning the protection of sensitive training data. Its defensive implications are profound, offering a robust mechanism for data owners to confidently share their valuable datasets with AI developers without the pervasive fear of unauthorized leakage.

For data owners, DLBox provides an unprecedented level of control and assurance. By encapsulating the training process within a confidential computing environment and enforcing strict rules about how data can be used, data owners can ensure that their data is utilized solely for benign model training. This capability is transformative for industries like healthcare, finance, and defense, where data privacy and compliance are paramount. Hospitals, for instance, can now contribute medical imaging data to AI research initiatives, fostering advancements in diagnostics and treatment, while being confident that patient data will not be exfiltrated or misused. DLBox effectively transforms the "borrowing" of data into a secure, auditable process, safeguarding the rights and intellectual property of data owners.

For AI developers and researchers, DLBox offers a pathway to access richer, more diverse datasets that were previously inaccessible due to security concerns. This can accelerate innovation and lead to the development of more accurate and robust AI models. While the framework imposes certain constraints on the training process (enforcing DGM rules), these constraints are designed to align with standard, ethical deep learning practices. Developers can focus on model architecture and optimization, trusting that the underlying data protection mechanism is robust.

Furthermore, DLBox's approach to defining "benign model training" through the DGM rules provides a clear, actionable framework for ethical AI development. It shifts the paradigm from an unconstrained environment to one where computational integrity is enforced. This can contribute to building trust in AI systems by ensuring that the models are trained in a transparently secure manner, reducing the risk of models being inadvertently or maliciously biased by specific data points that could compromise their fairness or robustness. The minimal 4% performance overhead also makes DLBox a practical solution that can be integrated into existing machine learning workflows without significantly impeding development cycles. Ultimately, DLBox empowers a more collaborative and secure future for data-driven AI, fostering innovation while rigorously upholding data privacy and ownership rights.

Key Takeaways

  • Critical Need for Data Protection: There is a significant and unresolved problem of protecting sensitive training data from leakage when shared with AI developers, hindering collaboration and innovation.
  • DLBox's Dual Approach: The framework combines a D-Sandbox (secure execution environment leveraging confidential computing like AMD SEV-SNP) to prevent direct data exfiltration, and a D-Oracle to monitor and enforce benign training practices.
  • Defining Benign Training with DGM Rules: DLBox introduces the DGM rules (Data augmentation uniformity, Gradient computation equality, Model update aggregation) as a novel way to define and enforce ethical, non-malicious model training.
  • Proven Leakage Prevention: Evaluations show DLBox effectively prevents data leakage, reducing it from 100% in baseline scenarios to almost zero, even against advanced embedding and biasing attacks.
  • Low Performance Overhead: DLBox achieves strong security guarantees with a modest average performance overhead of only 4% in training time for both image and language models, making it practical for real-world applications.
  • Fostering Secure Data Sharing: DLBox empowers data owners to share sensitive datasets with confidence, accelerating AI development in privacy-critical domains like healthcare by protecting their rights and intellectual property.

About the Speaker(s)

Jaewon Hur is a Postdoc at Georgia Tech. His research focuses on developing secure and robust systems, particularly in the context of deep learning and data privacy. In this talk, he presented DLBox, a framework designed to protect training data during AI model development, reflecting his expertise in confidential computing and secure machine learning. He collaborated on this work with Jan Eumong Kim, Yongi, and his post-advisor Pyongyang.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid systems security research that attacks a real and underappreciated problem: the trust gap between data owners and ML developers. The DGM rule formalization is a clean conceptual contribution, and the AMD SEV-SNP + proxy variable architecture is a credible engineering answer, not a hand-wavy 'just use TEEs' non-solution. The 4% overhead claim, if it holds up under scrutiny, makes this actually deployable rather than just academically interesting.

Heather Calloway (CISO) — WEAK

Technically credible systems research solving a real problem in the ML supply chain, but the talk stays in the lab and never reaches the institution. There is no path from 'postdoc at Georgia Tech' to 'CISO at a hospital system decides to adopt this,' and the paper doesn't try to build one.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025