SafeSplit: A Novel Defense Against Client-Side Backdoor Attacks in Split Learning

Phillip Rieger (Todd Damshot)

Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · Federated Learning 1

Overview

This talk introduces SafeSplit, a novel defense mechanism designed to protect split learning (SL) systems from client-side backdoor attacks. Given by Phillip Rieger from Todd Damshot (likely TU Darmstadt), the presentation highlights the unique vulnerabilities of split learning, particularly its sequential training paradigm, which renders many existing defenses from federated learning (FL) ineffective. Split learning is a critical collaborative learning scheme that allows multiple clients to jointly train large, complex deep neural networks (DNNs) even when individual clients lack the computational resources to handle the entire model, all while keeping their sensitive training data local.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction to Split Learning and backdoor vulnerability
  2. 1:20 Understanding forward and backward propagation in Split Learning
  3. 2:30 U-shaped Split Learning for enhanced privacy
  4. 4:40 Why backdoor attacks are challenging in Split Learning
  5. 5:40 Introducing SafeSplit: A novel defense mechanism
  6. 6:50 SafeSplit's benign checkpoint and rollback mechanism
  7. 7:30 Detailed explanation of SafeSplit's dynamic analysis technique

SafeSplit: A Novel Defense Against Client-Side Backdoor Attacks in Split Learning

Speakers: Phillip Rieger (Todd Damshot)

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=eMbVkGZo-bI

Overview

This talk introduces SafeSplit, a novel defense mechanism designed to protect split learning (SL) systems from client-side backdoor attacks. Given by Phillip Rieger from Todd Damshot (likely TU Darmstadt), the presentation highlights the unique vulnerabilities of split learning, particularly its sequential training paradigm, which renders many existing defenses from federated learning (FL) ineffective. Split learning is a critical collaborative learning scheme that allows multiple clients to jointly train large, complex deep neural networks (DNNs) even when individual clients lack the computational resources to handle the entire model, all while keeping their sensitive training data local.

The problem addressed is significant: while split learning offers substantial benefits in terms of privacy and resource efficiency, its distributed and sequential nature makes it highly susceptible to malicious clients injecting backdoors. Such backdoors can force the model to make specific, incorrect predictions when a predefined trigger is present in the input, without affecting benign predictions. SafeSplit proposes a "circle" architectural model for training, combined with an ensemble of static and dynamic analysis techniques, including a novel rotational frequency analyzer, to identify and bypass poisoned model updates before they can compromise the global model. This work is crucial for enabling the secure deployment of split learning in sensitive edge computing and privacy-preserving AI applications.

Background

▶ Watch: Introduction to Split Learning and backdoor vulnerability (0:00)

Collaborative machine learning, particularly with deep neural networks, has seen significant advancements, with federated learning (FL) emerging as a prominent paradigm. In FL, multiple clients each hold private datasets and collaboratively train a DNN. Critically, clients train their own local models and only share model parameters (weights) with a central server for aggregation, never their raw data. This preserves data privacy.

However, FL has a limitation: each client must be capable of training the entire DNN locally. This assumption breaks down when models become excessively large or clients operate on resource-constrained edge devices. This is where split learning (SL) offers a compelling alternative. In SL, the DNN itself is split into multiple parts. Typically, a "head" portion is deployed on the client side, and a "backbone" portion resides on a central server. During training, clients perform forward propagation up to a designated "cutting layer," sending the intermediate output (known as smash data) to the server. The server then completes the forward pass, calculates the loss, and initiates backward propagation, sending gradients back to the client.

A variant, U-shaped split learning, enhances privacy further by deploying a "tail" portion of the DNN back on the client side. In this configuration, the server only processes the backbone and never sees the raw input data or the final labels. While this significantly boosts privacy by keeping sensitive information entirely on the client side, it introduces a critical security vulnerability: clients are now responsible for the loss calculation. This power allows a malicious client to manipulate the loss function, facilitating the injection of backdoor attacks.

A backdoor attack aims to embed a hidden malicious behavior into a model. Specifically, for inputs containing a specific, often inconspicuous, trigger pattern (e.g., a pixel pattern in an image, a semantic feature like green cars), the model will consistently predict an adversary-chosen target label (e.g., "bird"). For all other benign inputs without the trigger, the model performs normally to avoid detection.

The inherent sequential nature of split learning training makes it uniquely vulnerable to such attacks. In federated learning, multiple clients contribute model updates in each round, allowing the server to compare these updates, identify outliers, and potentially exclude malicious contributions. In contrast, split learning often involves clients training sequentially. At any given step, only a single client is actively training a portion of the model. This means there is no immediate comparison base for detecting anomalous updates. If a malicious client injects a backdoor, the subsequent benign clients will receive and continue training on this poisoned model, propagating the malicious artifacts throughout the learning process. Detecting such poisoning after the fact is extremely challenging, and even if detected, it necessitates costly retraining, potentially from scratch. This fundamental architectural difference underscores why existing federated learning defenses are largely ineffective against client-side backdoors in split learning.

Key Findings

▶ Watch: U-shaped Split Learning for enhanced privacy (2:30)

The central finding of this research is the development of SafeSplit, a novel defense architecture specifically tailored to mitigate client-side backdoor attacks in split learning. The authors found that the sequential training process in split learning, unlike the round-based aggregation in federated learning, creates a unique vulnerability where a single malicious client can poison the model without immediate detection or comparison. Existing federated learning defenses, which often rely on outlier detection among multiple parallel updates, fail in this sequential context.

SafeSplit's key contributions and findings include:

  • A "Circle" Training Model: Instead of a traditional round-based model, SafeSplit conceptualizes the training process as a continuous "circle" where clients train sequentially. This allows for proactive detection and bypass of poisoned models before subsequent benign clients are affected.
  • Ensemble of Static and Dynamic Analyzers: SafeSplit combines two distinct analytical perspectives to comprehensively identify malicious model updates. This dual approach enhances robustness against sophisticated and adaptive attackers.
  • Novel Rotational Distance Metric: For dynamic analysis, SafeSplit introduces a unique rotational frequency analyzer that examines the orientation and rotation of weight matrices over time. This metric provides deeper insights into the training dynamics than traditional frequency analysis alone.
  • Effective Mitigation without Retraining: By identifying and rolling back to a benign checkpoint, SafeSplit can effectively bypass poisoned training efforts, preventing damage to the global model without requiring extensive and costly retraining.
  • Robustness against Adaptive Attacks: The combination of static and dynamic analysis, along with the non-differentiable nature of some of its components, makes SafeSplit resilient to various adaptive attack strategies, including those attempting to manipulate loss functions or blend malicious updates with benign data.

Technical Deep Dive

▶ Watch: Why backdoor attacks are challenging in Split Learning (4:40)

SafeSplit's core innovation lies in its proactive defense mechanism, which models the split learning training process not as discrete rounds but as a continuous "circle." This conceptual shift is crucial for addressing the sequential vulnerability. The primary goal is to ensure that before any client begins training a new model segment (e.g., W4), the incoming model checkpoint is verified as benign.

The defense operates by analyzing the backbone updates received from the clients. It employs a two-pronged approach, combining both static and dynamic analysis techniques to scrutinize these updates:

  1. Dynamic Analysis: Rotational Frequency Analyzers

This is a novel technique introduced by SafeSplit. It focuses on detecting subtle, malicious changes in the orientation and rotation of the neural network's weight matrices, particularly within the server-side backbone.

  • Weight Matrix Analysis: For a given weight matrix in the backbone, the system first calculates the row-wise and column-wise means.
  • Change Detection: These means are then multiplied with the original weight matrix to determine the X and Y changes in the matrix. This process helps to quantify how individual weights are shifting.
  • Angle and Rotation: From these X and Y changes, the angle of the weight matrix can be calculated, providing insight into its orientation.
  • Velocity and Frequency: To understand the dynamics of these changes, the system divides the angle by the time elapsed between updates, yielding a "velocity" of rotation. This velocity is then converted into a frequency representation. The underlying assumption is that benign training leads to gradual, predictable changes in weight matrices, while backdoor injection might cause abrupt or uncharacteristic rotations.
  • Anomaly Scoring: To identify poisoned models, SafeSplit calculates an anomaly score for each model update. This is done by first determining the "neighborhood" of a model update in the feature space derived from its rotational frequency analysis. The anomaly score is then the sum of the distances from that model to all other models within its neighborhood.
  • Benign Threshold: A model update is classified as benign if its anomaly score falls among the smallest 51% of all observed anomaly scores. This threshold identifies updates that exhibit consistent, low-deviation rotational behavior.
  1. Static Analysis: Frequency Domain Analyzers

This technique analyzes the model's parameters directly in the frequency domain, looking for patterns indicative of malicious alterations.

  • Frequency Transformation: The model's parameters (weights) are transformed into the frequency domain.
  • Component Observation: The authors observe that in early training epochs, low-frequency components of the model change significantly as new behaviors are learned. However, as the model approaches convergence in later phases, the ratio between low and high-frequency components stabilizes, and changes become less pronounced. Backdoor injection, especially in later stages, might disrupt this natural progression, causing unusual shifts in specific frequency components.
  • Anomaly Scoring: Similar to the dynamic analysis, SafeSplit isolates the lower frequency components and then calculates neighborhood distances to derive an anomaly score. Models with unusually high scores in these components, deviating from expected benign behavior, are flagged.

The Circle Architecture in Action:

When a client (e.g., Client 4) is about to train, SafeSplit performs the following checks on the latest received model update (e.g., from Client 3):

  1. The backbone update from Client 3 is subjected to both the dynamic (rotational frequency) and static (frequency domain) analyses.
  2. If both analyses classify the update as benign, then Client 4 proceeds to train based on this model.
  3. If either analysis flags the update as potentially malicious (e.g., Client 3's update is poisoned), SafeSplit does not use it. Instead, it rolls back to the previous model checkpoint (e.g., from Client 2) and repeats the analysis.
  4. This rollback and re-evaluation process continues until a model checkpoint that is deemed benign by both metrics is found. This ensures that subsequent benign clients always start training from a clean, uncompromised state, effectively bypassing any poisoned contributions without needing to retrain the entire model or identify the specific malicious client.

The ensemble approach of combining static and dynamic analysis is critical. An adversary attempting to bypass one defense might inadvertently trigger the other. For instance, an attack designed to be stealthy in the frequency domain might still cause detectable rotational anomalies in the weight matrices, and vice-versa. This redundancy significantly increases the difficulty for an attacker to adaptively circumvent the defense.

Demo / Proof of Concept

▶ Watch: SafeSplit's benign checkpoint and rollback mechanism (6:50)

While the talk does not describe a live demonstration or a specific step-by-step proof of concept during the presentation, the speaker clearly states that SafeSplit was thoroughly evaluated across multiple datasets and various model architectures. This evaluation included comparisons against adaptations of existing defenses originally developed for federated learning. As anticipated, these federated learning-based defenses proved largely ineffective in the sequential training environment of split learning.

The evaluation specifically considered different types of backdoor attacks:

  • Semantic triggers: For example, making green cars be mispredicted as "birds."
  • Pixel triggers: Specific pixel patterns, suitable for datasets like MNIST.
  • Sophisticated attack strategies: This included scenarios where adversaries attempted to manipulate the loss function to adapt to the defense or mixed benign data with malicious updates to make the attack less conspicuous. The talk emphasized that the ensemble of SafeSplit's static and dynamic techniques successfully prevented such adaptive attacks from succeeding.

The results, as referenced in the paper, demonstrate SafeSplit's capability to effectively mitigate these diverse backdoor threats while preserving the integrity of the benign training process.

Defensive Implications

▶ Watch: Detailed explanation of SafeSplit's dynamic analysis technique (7:30)

The introduction of SafeSplit carries significant implications for securing collaborative machine learning deployments, particularly those leveraging split learning.

  1. Mandatory for U-shaped Split Learning: For any system employing U-shaped split learning, where clients handle loss calculation and thus have greater control over model updates, integrating a robust defense like SafeSplit becomes paramount. The privacy benefits of U-shaped SL are substantial, but they must be balanced with strong security measures to prevent client-side manipulation.
  2. Rethinking Training Paradigms: SafeSplit's "circle" architecture challenges the conventional round-based thinking in distributed learning. Defenders should consider adopting continuous monitoring and checkpoint verification mechanisms, especially in environments where sequential updates are common. This proactive approach prevents the propagation of malicious artifacts rather than attempting to remediate a poisoned model after the fact.
  3. Enhanced Monitoring Metrics: The novel rotational frequency analysis provides new, powerful metrics for monitoring model updates. Beyond traditional statistical analyses, tracking changes in weight matrix orientation and velocity can offer a more granular and sensitive indicator of anomalous behavior, which could be integrated into broader MLOps security pipelines.
  4. Resilience Against Adaptive Adversaries: The ensemble nature of SafeSplit's defense (combining static and dynamic analysis) is a crucial lesson. Security architects should aim for multi-layered defenses that are non-differentiable where possible. This makes it significantly harder for sophisticated attackers to craft adaptive strategies that can bypass the defense by calculating gradients against it. Even heuristic attempts by attackers might leave detectable "subtle differences" that an ensemble defense can catch.
  5. Impact on Resource-Constrained Environments: By enabling secure split learning, SafeSplit helps unlock the potential of deploying large, complex DNNs in edge scenarios and on resource-constrained devices without compromising security. This is vital for applications like medical imaging, IoT analytics, and autonomous systems where data privacy and model integrity are critical.
  6. Addressing Non-IID Data: The Q&A session highlighted the challenge of non-IID (non-independently and identically distributed) data, a common issue in distributed learning. SafeSplit's approach of looking for consistency and inversion of training efforts, rather than just "difference," is key. Defenders need to understand that benign clients with highly divergent data might naturally produce "outlier-like" updates. The defense must distinguish between benign deviations and malicious intent to invert or corrupt the model's learning trajectory.

In summary, SafeSplit provides a concrete framework for mitigating a significant threat in split learning, pushing the boundaries of secure collaborative AI and offering practical guidance for developers and security professionals in this rapidly evolving field.

Key Takeaways

  • Split Learning's Unique Vulnerability: Unlike federated learning, split learning's sequential training paradigm makes it highly susceptible to client-side backdoor attacks, as there's no comparison base for detecting malicious updates, and poisoned models can propagate easily.
  • Proactive "Circle" Architecture: SafeSplit introduces a novel "circle" training model that enables proactive detection and bypassing of poisoned model updates before they can affect subsequent benign clients, preventing widespread model corruption.
  • Ensemble Detection for Robustness: The defense combines two distinct analysis techniques – static frequency domain analysis and a novel dynamic rotational frequency analysis – to provide comprehensive and robust detection against various backdoor types and adaptive attack strategies.
  • Rotational Frequency Analysis: A key innovation is the rotational frequency analyzer, which dynamically examines changes in the orientation and rotation velocity of weight matrices in the model's backbone, offering a sensitive indicator of malicious manipulation.
  • Effective Mitigation without Retraining: SafeSplit can effectively identify and roll back to benign model checkpoints, allowing the training process to bypass poisoned contributions without the need for costly and time-consuming full model retraining.
  • Resilience Against Adaptive Adversaries: The combination of non-differentiable components and an ensemble of detection techniques makes SafeSplit highly resistant to adaptive attacks, which struggle to circumvent both analysis methods simultaneously.

About the Speaker(s)

Phillip Rieger is a researcher affiliated with Todd Damshot, where he and his collaborators developed SafeSplit. His work focuses on addressing security vulnerabilities, particularly backdoor attacks, in collaborative machine learning paradigms like split learning. The specific affiliation "Todd Damshot" is likely a reference to TU Darmstadt, a prominent technical university known for its research in computer science and security.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Legitimate academic security research on a real and underexplored problem — backdoor attacks in split learning — with a technically coherent defense mechanism. The rotational frequency analyzer is a genuinely novel metric, and the ensemble approach is sound in principle, but this is a conference paper presentation, not a practitioner talk, and the gap between the academic contribution and real-world deployability is wide enough to matter.

Heather Calloway (CISO) — WEAK

Technically credible research on a real vulnerability in split learning, but it never crosses into defender territory. No operator, executive, or security leader leaves knowing what to do — only that a new defense mechanism exists in a lab.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025