Try to Poison My Deep Learning Data? Nowhere to Hide Your Trajectory Spectrum!
Yansong Gao
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · ML Backdoors
Overview
In the rapidly evolving landscape of deep learning, the quality and integrity of training data are paramount, especially for the development of sophisticated models like large language models and other foundation models. However, the acquisition of high-quality, diverse datasets is a significant challenge, often leading organizations to adopt a data-as-a-service (DaaS) business model. This model typically involves data curators outsourcing data collection and annotation to numerous freelancers or data contributors, who may not always be fully trusted. This talk by Yansong Gao presents a critical investigation into the vulnerabilities arising from such untrusted data sources, specifically focusing on data poisoning attacks designed to inject backdoors into deep learning models.
Key moments
- 1:00 The data poisoning problem in data-as-a-service
- 3:30 Limitations of existing data cleansing methods and requirements
- 5:30 Initial insight: Visually distinguishing benign and poison loss trajectories
- 6:10 Challenge: Overlapping loss trajectories in t-SNE
- 6:25 Solution: Spectral transformation for enhanced separation
- 7:00 Refinement: Truncating loss trajectory after convergence
- 8:10 Overview of the complete data cleansing pipeline
- 8:25 Rationale for choosing DBScan over K-means clustering
Try to Poison My Deep Learning Data? Nowhere to Hide Your Trajectory Spectrum!
Speakers: Yansong Gao
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=mOvzHOc1-30
Overview
In the rapidly evolving landscape of deep learning, the quality and integrity of training data are paramount, especially for the development of sophisticated models like large language models and other foundation models. However, the acquisition of high-quality, diverse datasets is a significant challenge, often leading organizations to adopt a data-as-a-service (DaaS) business model. This model typically involves data curators outsourcing data collection and annotation to numerous freelancers or data contributors, who may not always be fully trusted. This talk by Yansong Gao presents a critical investigation into the vulnerabilities arising from such untrusted data sources, specifically focusing on data poisoning attacks designed to inject backdoors into deep learning models.
The core problem addressed is the responsibility of data curators to cleanse potentially poisoned data before it is delivered to model providers. Existing defenses against backdoors are often ill-suited for this DaaS paradigm, as they typically require intervention from the model provider or rely on assumptions (like access to clean data) that are not practical for a curator. The research introduces a novel data cleansing methodology that leverages the loss trajectory of training samples and spectral transformation to effectively identify and isolate poisoned data, even under challenging conditions such as extremely low poison rates, diverse attack types, and without any prior knowledge of clean data. This work offers a crucial advancement for securing the deep learning supply chain and ensuring the trustworthiness of outsourced datasets.
Background
▶ Watch: The data poisoning problem in data-as-a-service (1:00)
The proliferation of deep learning applications has created an immense demand for vast, high-quality datasets. To meet this demand, a common industry practice is the data-as-a-service (DaaS) model. In this setup, data curators act as intermediaries, collecting or commissioning data from a multitude of data contributors (often freelance workers, as exemplified by platforms like Appen, Amazon Mechanical Turk, Clickworker with its six million freelancers, and Telus). These curators then sell or license this aggregated and annotated data to model providers who use it to train their deep learning models.
A significant security vulnerability arises in this ecosystem: the untrusted data contributor. If a contributor is malicious, they can introduce poisoned samples into the dataset. These poisoned samples are specifically crafted to inject backdoors into the models trained on them. A backdoored model will behave normally on benign inputs but will exhibit a malicious, predetermined behavior when presented with a specific trigger (e.g., a small patch, an imperceptible perturbation). For example, a model trained with a backdoor might misclassify a stop sign as a yield sign if a particular trigger is present on the sign.
In the DaaS model, the onus of ensuring data integrity falls squarely on the data curator. It is their responsibility to cleanse the data of any poisoning artifacts before it is distributed to model providers. However, existing defenses against backdoors largely fall into three categories:
- Prevision-based defenses: These methods aim to prevent backdoors from being injected in the first place. They are typically the responsibility of the model provider or require control over the training environment, making them inapplicable for a data curator.
- Model-based detection: These techniques analyze the trained model to detect backdoor presence. Again, this is a post-training activity and thus falls under the model provider's domain, not the data curator's pre-delivery cleansing task.
- Database detection: These methods focus on identifying and removing poisoned samples directly from the dataset. This is the only category suitable for data curators. Within database detection, approaches performed during the inference phase are also unsuitable, as the curator needs to act before training. Therefore, only training-phase database detection that can be performed as a one-time operation is viable for the DaaS scenario.
Beyond this fundamental requirement, the research identifies five crucial practical requirements for any effective data cleansing method in a DaaS context:
- Agnostic to modality: The method must work irrespective of the data type (e.g., images, audio, text), as curators handle diverse data streams.
- Agnostic to task: It must be effective across different deep learning tasks (e.g., classification, regression, object detection, segmentation), whereas many existing methods are confined to classification.
- Insensitive to poison rate: The method must reliably detect backdoors even when the percentage of poisoned samples is extremely low. Attacks like SORT (CCS 2023) have demonstrated successful backdoor injection with poison rates as low as 0.25% (25 samples out of 50,000).
- No prior knowledge of attack methodology: The curator cannot be expected to know the specific trigger design or backdoor type used by an attacker.
- No clean data access: This is perhaps the most challenging requirement. In a real-world DaaS scenario, the curator may not have access to a clean, unpoisoned subset of data to use as a baseline or reference.
A review of existing representative data cleansing methods reveals that most fail to meet several of these requirements. They often rely on access to clean data, are sensitive to specific trigger designs or backdoor types, and are typically limited to classification tasks. This starkly highlights the need for a novel approach capable of addressing these practical constraints.
Key Findings
▶ Watch: Initial insight: Visually distinguishing benign and poison loss trajectories (5:30)
The research introduces a groundbreaking approach that addresses the limitations of existing data cleansing methods by leveraging fundamental characteristics of how deep learning models learn from data. The key findings are rooted in several interconnected insights:
1. The Loss Trajectory as a Universal Indicator:
The first crucial insight is that the loss trajectory—the sequence of loss values a sample experiences during the training process—can serve as a powerful indicator to distinguish between poisoned and benign samples. The authors observed that, for many common backdoor attacks (including source-agnostic or universal backdoors and source-specific or partial backdoors), poisoned samples often exhibit a lower loss than benign samples, particularly in the early epochs of training. This phenomenon occurs because poisoned samples, designed to activate a specific malicious behavior, might be "easier" for the model to learn or fit to the target malicious label early on. However, the talk also highlights an important counter-example: advanced attacks like the SORT attack (CCS 2023), which operate with extremely low poison rates (e.g., 25 samples out of 50,000), can actually result in higher loss for poisoned samples compared to benign ones. Despite these variations, the critical observation is that the loss trajectories of benign and poisoned samples are visually distinguishable, suggesting an underlying difference that can be programmatically exploited.
2. Loss as a Modality and Task-Agnostic Metric:
A significant advantage of using the loss trajectory is its universality. Loss is an intrinsic metric in virtually all deep learning training processes, regardless of the data modality (e.g., images, audio, text) or the specific learning task (e.g., classification, regression, object detection, segmentation). This universality directly addresses two of the critical practical requirements for data curators: modality agnosticism and task agnosticism. By focusing on loss, the proposed method bypasses the need for specialized techniques tailored to different data types or tasks, offering a truly generalizable solution.
3. Spectral Transformation for Amplifying Differences:
While loss trajectories show visual differences, a direct clustering of their raw embeddings (e.g., using t-SNE) reveals significant overlap between benign and poisoned samples, making clear separation difficult. The third key insight involves treating the loss trajectory as a time series signal. By applying spectral transformation (e.g., converting the signal into the frequency domain), subtle differences in the temporal patterns of loss can be significantly amplified. This transformation helps to expose distinguishing features that are not readily apparent in the raw time-domain signal. The t-SNE visualizations presented in the talk vividly illustrate this: raw trajectory embeddings are highly intermingled, but after spectral transformation, the poisoned samples show much better separation from benign samples.
4. Truncation for Enhanced Salience:
Further refining the approach, the researchers found that truncating the loss trajectory to focus on the period after convergence yielded even better results. The relative differences between benign and poisoned samples become more salient and pronounced once the model has largely converged. Combining this truncation with spectral transformation leads to near-perfect separation of poisoned samples from benign ones in the spectral domain. This combination effectively addresses the challenge of distinguishing subtle differences, particularly critical for detecting low-poison-rate attacks.
In summary, these key findings demonstrate that by intelligently analyzing the loss trajectory—treating it as a universal time series signal, truncating it, and applying spectral transformation—it is possible to robustly identify poisoned samples. This foundational understanding forms the basis of a data cleansing methodology that effectively tackles the challenging requirements of the data-as-a-service model, including operating without clean data access or prior knowledge of attack specifics.
Technical Deep Dive
▶ Watch: Solution: Spectral transformation for enhanced separation (6:25)
The proposed methodology, which we can refer to as Telltale (inferred from the paper title), is designed as a robust pipeline for detecting poisoned samples within a dataset, specifically tailored for the stringent requirements of data curators in a DaaS environment. The pipeline consists of four main steps:
- Loss Trajectory Collection and Truncation:
The first step involves collecting the loss trajectory for each individual sample during the deep learning model's training process. This means recording the loss value computed for each sample at various epochs. Critically, the method does not require any specialized training procedures; it simply monitors the loss during standard deep learning training.
Once the full loss trajectory is recorded, a truncation step is applied. Instead of using the entire trajectory from the very beginning of training, the method focuses on the loss values observed after the model has largely converged. This is because, as observed in the key findings, the relative differences between benign and poisoned samples become more pronounced and stable in the later stages of training. This truncation helps to filter out noise from the initial, unstable learning phases and highlight the more salient discriminatory patterns.
- LSTM-based Encoder for Dimension Reduction:
The truncated loss trajectories are still high-dimensional time series. To efficiently process these signals and extract meaningful features, an LSTM-based encoder is employed. A Long Short-Term Memory (LSTM) network is a type of recurrent neural network particularly adept at processing sequential data like time series. The encoder takes the high-dimensional loss trajectory as input and compresses it into a lower-dimensional, fixed-size embedding vector. This embedding captures the essential temporal characteristics and patterns of the loss trajectory, effectively performing dimension reduction while preserving critical information. This step is crucial for transforming the raw time series into a format suitable for subsequent analysis.
- Spectral Transformation:
Following dimension reduction, the embedding vectors (representing the loss trajectories) undergo a spectral transformation. While the transcript doesn't specify the exact transformation, the discussion of "frequency domain" implies techniques like the Fast Fourier Transform (FFT) or similar spectral analysis methods. The core idea is to convert the time-domain signal (the loss trajectory's temporal pattern) into the frequency domain. This transformation helps to amplify subtle, recurring patterns and differences that might be obscured in the time domain. By analyzing the frequency components, the method can highlight distinct periodicities or spectral signatures that differentiate poisoned samples from benign ones. The effectiveness of this step is visually demonstrated through t-SNE plots, where samples that were highly overlapped in raw embedding space become significantly separated after spectral transformation.
- Clustering with DB-scan:
The final step involves applying a clustering algorithm to the spectrally transformed embeddings to segment the dataset into distinct groups. The choice of clustering algorithm is critical here. The method utilizes DB-scan (Density-Based Spatial Clustering of Applications with Noise) instead of more common algorithms like K-means.
The primary reason for choosing DB-scan is its ability to operate without requiring prior knowledge of the number of clusters. In a real-world DaaS scenario, a data curator might receive a dataset that is entirely clean (meaning only one cluster of benign samples) or partially poisoned (meaning two clusters: one benign, one poisoned). K-means would require the curator to pre-specify k (the number of clusters), which is an impractical requirement without knowing if the data is poisoned. DB-scan, conversely, identifies clusters based on the density of data points and can naturally discover an arbitrary number of clusters, including identifying noise points. This makes it robust and practical for the "no prior knowledge" requirement. The DB-scan algorithm effectively separates the dense cluster(s) of benign samples from the potentially sparser or distinct cluster(s) formed by poisoned samples in the spectrally transformed space.
Extensive Evaluation and Comparison:
The efficacy of the Telltale methodology was validated through extensive evaluations across a wide range of attack scenarios and practical constraints:
- Backdoor Types: Both universal backdoors (source-agnostic, affecting any input with a trigger) and partial backdoors (source-specific, only activating on inputs from a specific source class with a trigger) were considered.
- Trigger Types: The method was tested against diverse trigger designs, including pad triggers, blended triggers, dynamic triggers, imperceptible triggers, and ISBA (Image-Specific Backdoor Attack).
- Modalities and Tasks: The universality of the loss metric was confirmed by evaluating the method on image, audio, and text data, and across classification and regression tasks.
- Poison Rates: Performance was consistently strong even with varying and extremely low poison rates, demonstrating insensitivity to this critical parameter.
The detection performance for universal backdoors, for example, reached up to 97%, with a very low false positive rate (rarely misclassifying benign samples as poisoned).
Crucially, Telltale was benchmarked against state-of-the-art backdoor detection methods, including SORT (CCS 2023) and CT (often a competitive baseline from recent security conferences like EuroS&P). Comparisons were made in three critical settings: universal trigger attacks, partial backdoor attacks, and, significantly, on benign (clean) datasets. The results consistently showed Telltale outperforming its competitors. A notable highlight was Telltale's robustness on benign datasets, where it correctly identified the absence of poisoning. In contrast, the SORT method, while effective in some scenarios, was found to be problematic on benign datasets, falsely recognizing a large fraction of clean samples as poisoned. This underscores Telltale's practical value, as data curators need a method that performs reliably even when the data is genuinely clean, a scenario often overlooked in research evaluations.
The authors emphasize that Telltale is the only methodology presented that simultaneously satisfies all the practical requirements for data cleansing in the DaaS model: modality-agnostic, task-agnostic, insensitive to poison rate, no prior knowledge of attack methodology, and no clean data access. The source code for the method has been released, enabling further research and adoption.
Demo / Proof of Concept
▶ Watch: Refinement: Truncating loss trajectory after convergence (7:00)
While the talk focuses heavily on the methodology and extensive evaluation results, it does not detail a specific live demonstration or a step-by-step proof-of-concept setup in the traditional sense. Instead, the "demonstration" of the method's efficacy is primarily conveyed through the comprehensive experimental results presented. The evaluations cover a broad spectrum of attack types, trigger designs, data modalities (image, audio, text), and tasks (classification, regression), along with comparisons against state-of-the-art baselines.
The visualizations, such as t-SNE plots showing the separation of benign and poisoned samples after spectral transformation, serve as a powerful proof of concept for the underlying insights. The high detection performance (up to 97%) and low false positive rates across various challenging conditions effectively demonstrate that the proposed approach is not just theoretically sound but practically effective in identifying poisoned data without needing clean data access or prior attack knowledge. The release of the source code further enables researchers and practitioners to replicate and verify the method's capabilities.
Defensive Implications
▶ Watch: Rationale for choosing DBScan over K-means clustering (8:25)
The Telltale methodology offers profound defensive implications for the deep learning ecosystem, particularly for entities operating within the data-as-a-service (DaaS) model.
For Data Curators:
The most direct beneficiaries are data curators. This method provides them with a robust, automated, and practical tool to fulfill their critical responsibility of ensuring data integrity. By implementing Telltale, curators can:
- Enhance Trustworthiness: Offer model providers datasets that have been verifiably cleansed of common backdoor poisoning attacks, significantly increasing the value and trustworthiness of their service.
- Meet Practical Requirements: Leverage a solution that is agnostic to data modality and task, insensitive to low poison rates, requires no prior knowledge of attack specifics, and crucially, operates without needing access to clean data. This addresses the exact pain points of real-world DaaS operations.
- Automate Cleansing: Integrate the pipeline into their data processing workflows, enabling efficient and scalable detection and removal of poisoned samples as a one-time operation before data distribution.
- Reduce Liability: Mitigate the risk of inadvertently supplying poisoned data that could lead to compromised downstream models and potentially severe consequences for their clients.
For Model Providers:
While the method operates at the data curator's end, model providers indirectly benefit significantly. They gain:
- Higher Data Confidence: Increased assurance that the data they purchase or subscribe to from curators is free from backdoors, reducing the risk of training vulnerable models.
- Reduced Attack Surface: Less need to implement their own complex and often resource-intensive backdoor detection mechanisms, as the data is pre-cleansed.
- Improved Model Reliability: Models trained on cleaner data are inherently more robust and less susceptible to adversarial manipulation at inference time.
For the Deep Learning Community:
More broadly, this research contributes to a more secure and resilient deep learning supply chain.
- Standardization of Defenses: It pushes towards more standardized and effective data integrity checks in outsourced data scenarios, which are becoming increasingly common.
- Inspiration for Future Research: The success of leveraging loss trajectories and spectral transformation as universal indicators may inspire novel defensive strategies against other forms of data-centric attacks or for general data quality assessment in deep learning.
- Awareness of Overlooked Scenarios: The emphasis on testing against benign datasets and the challenges posed by extremely low poison rates highlight critical aspects often overlooked in security evaluations, promoting more rigorous and realistic benchmarking in the field.
In essence, Telltale empowers data curators to act as a crucial security gatekeeper, safeguarding the integrity of data that fuels the next generation of AI systems and fostering greater trust in the deep learning ecosystem.
Key Takeaways
- Data poisoning in Data-as-a-Service (DaaS) is a critical threat: Untrusted data contributors can inject backdoors into datasets, necessitating robust cleansing by data curators before data reaches model providers.
- Existing defenses are largely unsuitable for DaaS: Most current backdoor detection methods fail to meet practical requirements for data curators, such as modality/task agnosticism, insensitivity to low poison rates, no prior attack knowledge, and especially, no access to clean data.
- Loss trajectory is a universal, insightful metric: The dynamic behavior of sample loss during training (loss trajectory) provides a powerful, modality- and task-agnostic signal for distinguishing poisoned samples from benign ones.
- Spectral transformation unlocks hidden differences: Treating loss trajectories as time series and applying spectral transformation (e.g., in the frequency domain) dramatically amplifies subtle differences, enabling effective separation of poisoned samples. Truncating trajectories after convergence further enhances this effect.
- Robust pipeline for practical deployment: The proposed method utilizes an LSTM-based encoder for dimension reduction, followed by spectral transformation, and then DB-scan clustering. DB-scan is crucial as it doesn't require prior knowledge of the number of clusters, making it suitable for datasets that might be either clean or poisoned.
- Superior performance across diverse challenges: Extensive evaluations demonstrate the method's effectiveness against various backdoor types, trigger designs, data modalities (image, audio, text), and tasks (classification, regression), consistently outperforming state-of-the-art baselines, particularly in its robustness on benign datasets.
About the Speaker(s)
The talk was presented by Yansong Gao. As indicated in the opening remarks, this was a "joint work," suggesting collaboration with other researchers. The transcript does not provide specific details about Yansong Gao's title, affiliation, or other biographical information beyond their role as a presenter for this research. Based on the context of the NDSS Symposium, it is common for speakers to be researchers or academics in the field of cybersecurity.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic security research with a clear problem framing and a technically coherent solution — loss trajectory + spectral transformation + DBSCAN is a sensible and novel pipeline for the DaaS data-cleansing scenario. The contribution is real, but this is a conference paper presentation, not a security talk, and the gap between 'academically valid' and 'operationally moves the needle' is wide enough to matter.
Heather Calloway (CISO) — WEAK
Technically credible research on detecting poisoned training data in outsourced data pipelines, with a genuinely practical framing around the DaaS model. But this is a pure research presentation — it ends at the algorithm and never reaches the institutional, contractual, or operational questions that would make it actionable for anyone running a security program.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025