URVFL: Undetectable Data Reconstruction Attack on Vertical Federated Learning
Duanyi Yao
Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · Federated Learning 1
Overview
This talk introduces URVFL (Undetectable Data Reconstruction Attack on Vertical Federated Learning), a novel and potent data reconstruction attack designed to operate stealthily against Vertical Federated Learning (VFL) systems. Presented by Duanyi Yao, this research, a collaborative effort with Suni Shangongo and Ging Pan, addresses a critical vulnerability in VFL: the ability of a malicious active client to reconstruct private data features from passive clients, even when robust detection mechanisms are in place. The core innovation lies in its capacity to generate malicious gradients that are indistinguishable from benign ones, thereby evading state-of-the-art detection strategies.
Key moments
- 0:00 Introduction to VFL and data reconstruction attacks
- 2:00 Threat model and challenges for malicious VFL attacks
- 3:20 URVFL: proposed undetectable data reconstruction attack method
- 4:40 Key innovation: Discriminator with Auxiliary Classifier (DSA)
- 6:00 Benefits of DSA for improved embedding transfer performance
- 8:00 Experimental results: URVFL's superior reconstruction performance
- 9:20 Crucial finding: URVFL maintains performance under detection
URVFL: Undetectable Data Reconstruction Attack on Vertical Federated Learning
Speakers: Duanyi Yao
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=TYPzIH1VpyE
Overview
This talk introduces URVFL (Undetectable Data Reconstruction Attack on Vertical Federated Learning), a novel and potent data reconstruction attack designed to operate stealthily against Vertical Federated Learning (VFL) systems. Presented by Duanyi Yao, this research, a collaborative effort with Suni Shangongo and Ging Pan, addresses a critical vulnerability in VFL: the ability of a malicious active client to reconstruct private data features from passive clients, even when robust detection mechanisms are in place. The core innovation lies in its capacity to generate malicious gradients that are indistinguishable from benign ones, thereby evading state-of-the-art detection strategies.
Vertical Federated Learning is a collaborative machine learning paradigm where multiple clients, each holding different features of the same dataset, jointly train a model without sharing their raw data. Privacy is theoretically preserved by exchanging only intermediate representations, known as embeddings. However, this work demonstrates that if the active client – the party with access to labels and a portion of features – is malicious, it can exploit these embeddings to infer or reconstruct the private data features of passive clients. The significance of URVFL stems from its focus on malicious adversaries who can deviate from standard protocols, a threat model often overlooked by prior work, and its unprecedented ability to remain undetected while achieving high-fidelity data reconstruction.
The talk highlights two major challenges in launching effective data reconstruction attacks against VFL: the heterogeneity and black-box nature of distributed features, and the detectability of malicious gradient manipulations. Previous malicious attacks were often easily identified by gradient comparison methods. URVFL directly tackles these challenges, providing a robust methodology that not only achieves superior reconstruction quality across diverse datasets but also consistently evades detection by established security measures like SweetGuard and Gradient Screener. This research underscores an urgent need for re-evaluating and enhancing defensive strategies in VFL to counteract such sophisticated, stealthy threats.
Background
▶ Watch: Introduction to VFL and data reconstruction attacks (0:00)
Vertical Federated Learning (VFL) is a specialized form of federated learning designed for scenarios where data is vertically partitioned, meaning different clients hold distinct feature sets for the same set of entities. In a typical VFL setup, there is an active client who possesses the labels and a subset of data features, and one or more passive clients, each holding other, complementary data features. During the training process, each client trains a bottom model locally to compute intermediate representations, or embeddings, from their private data features. These embeddings are then sent to the active client, who combines them with its own features and uses a top model for final prediction or classification. The fundamental privacy guarantee of VFL relies on the principle that only these embeddings, not the raw private data, are exchanged.
Despite this design, VFL is not immune to privacy attacks. Data reconstruction attacks represent a significant threat, where an adversary attempts to infer or reconstruct the original private data features from the shared embeddings. These attacks typically fall into two categories based on the adversary's capabilities:
- Honest-but-curious adversaries: These adversaries strictly adhere to the VFL training protocols but attempt to extract private information from the data they legitimately receive (e.g., embeddings, gradients). Many early works on VFL privacy focused on defending against this type of adversary.
- Malicious adversaries: These are more powerful adversaries who can actively manipulate or violate the training protocols to achieve their goals, such as stealing private data. While posing a greater threat, fewer previous works have thoroughly explored attacks under this stronger adversary model, partly due to the increased complexity of remaining undetected.
URVFL focuses on this more potent malicious adversary model. The threat model considered by the authors is specific:
- The adversary is the active client. This is a critical assumption, as the active client has central control over the top model and receives embeddings from all passive clients.
- The adversary is malicious, meaning it can deliberately alter or violate the VFL protocol to facilitate data reconstruction.
- The adversary has access to a small auxiliary dataset. This dataset shares a similar data distribution with the training data but is distinct from it, providing a crucial resource for pre-training attack components.
Launching effective data reconstruction attacks against VFL, especially under a malicious adversary model, presents two primary challenges that URVFL aims to overcome:
- Distributed Features and Limited Model Access: The target client's model (the passive client's bottom model) is often a black-box to the active adversary. Furthermore, the data feature space held by the passive client can be highly heterogeneous and distinct from the features the active client possesses. This makes it inherently difficult for the adversary to infer the structure or content of the remote data features based solely on the received embeddings.
- Detectability of Malicious Gradients: Malicious data reconstruction attacks often necessitate altering the VFL training task or manipulating the gradients exchanged during training. Such alterations, however, can easily be detected by comparing the manipulated gradients against what would be expected from an honest training process. Current detection methods, such as SweetGuard and Gradient Screener, are specifically designed to identify these discrepancies, effectively thwarting many malicious attacks by flagging anomalous gradient behavior. The challenge for a malicious adversary, therefore, is to manipulate gradients in a way that facilitates reconstruction without triggering these detection mechanisms.
URVFL directly addresses these formidable challenges by introducing a sophisticated attack methodology designed for stealth and efficacy, pushing the boundaries of what is considered secure in VFL.
Key Findings
▶ Watch: URVFL: proposed undetectable data reconstruction attack method (3:20)
The research presented on URVFL reveals several critical findings that significantly advance the understanding of privacy vulnerabilities in Vertical Federated Learning, particularly concerning malicious and undetectable attacks.
Firstly, URVFL demonstrates the feasibility and high efficacy of an undetectable data reconstruction attack against VFL systems. The attack consistently achieves superior reconstruction quality compared to existing state-of-the-art baselines, even when faced with robust detection mechanisms. For instance, on tabular datasets, URVFL achieved a reconstruction error of 0.42, significantly outperforming the suboptimal baseline FSG, which registered 0.5. Similarly, on image datasets like CIFAR-10, URVFL (specifically its variant URVFL-Sync) achieved a reconstruction error of 0.01, a marked improvement over FSG's 0.03. These quantitative results highlight URVFL's ability to extract high-fidelity private data.
Secondly, a core contribution and key finding is the development of the Discriminator with Auxiliary Classifier (DAC). This novel component is central to URVFL's ability to generate malicious gradients that are indistinguishable from honest gradients. The DAC mechanism effectively guides the target model to mimic the embedding distribution of the adversary's encoder while also incorporating label information. This strategic use of labels is crucial, as it not only improves the embedding transfer performance but critically makes the gradients less detectable by current methods. The speaker presented a visualization showing that embeddings generated with DAC are significantly more similar to original embeddings compared to those generated by a standard discriminator, evidenced by a lower cosine distance (e.g., 0.02 for DAC on MNIST vs. 0.04 for standard discriminator).
Thirdly, and perhaps most alarmingly, the research unequivocally shows that URVFL can consistently evade state-of-the-art detection strategies such as SweetGuard (SG) and Gradient Screener (GS). While other malicious attacks like AJN and FSG experienced significant performance drops when confronted with these detectors (e.g., PSNR dropping from 22 to 10), URVFL maintained consistent reconstruction quality. The detection scores for URVFL consistently remained above the detection threshold, indicating that the attack was not flagged. In contrast, other methods fell below the threshold, signifying their easy detection. This finding fundamentally challenges the efficacy of current VFL defense mechanisms against sophisticated, protocol-violating adversaries.
Finally, the visual evidence presented in the talk starkly illustrates URVFL's superiority. Reconstructed images from URVFL were clearly meaningful and recognizable, even under detection, whereas other methods failed to reconstruct any meaningful pixels when detection was active. This qualitative evidence reinforces the quantitative metrics, solidifying URVFL's position as a highly effective and stealthy data reconstruction attack, demanding immediate attention from the VFL security community.
Technical Deep Dive
▶ Watch: Key innovation: Discriminator with Auxiliary Classifier (DSA) (4:40)
URVFL's effectiveness stems from a meticulously designed three-step process: Pre-training, Malicious Gradient Generation, and Data Reconstruction. Each step plays a crucial role in enabling high-fidelity data reconstruction while simultaneously ensuring the attack remains undetectable.
1. Pre-training Phase
The initial phase of URVFL involves pre-training an encoder-decoder structure using the adversary's small auxiliary dataset. The primary goal here is two-fold:
- Encoder Training: The encoder is trained to produce embeddings that are structurally similar to those generated by the target model (the passive client's bottom model). This involves learning a mapping from raw data features to an embedding space.
- Decoder Training: Concurrently, a decoder is trained to effectively recover the original data features from these generated embeddings. This step is critical for the final reconstruction, as it teaches the decoder how to reverse the embedding process.
By pre-training on an auxiliary dataset that shares a similar distribution with the actual training data, the adversary establishes a foundational understanding of the feature-to-embedding and embedding-to-feature transformations. This prepares the decoder to perform accurate reconstructions once the malicious gradients allow for effective embedding transfer.
2. Malicious Gradient Generation (with DAC)
This is the most innovative and critical step of URVFL, where the adversary actively manipulates the VFL training process to transfer the embedding distribution from their pre-trained encoder to the target model of the passive client. The core innovation here is the Discriminator with Auxiliary Classifier (DAC).
Traditional approaches for embedding distribution transfer often rely on standard discriminators, which typically minimize the Jensen-Shannon (JS) distance between embedding distributions. However, as highlighted in the talk, this approach has two significant shortcomings:
- Ignores Label Information: Standard discriminators often overlook the crucial label information associated with the data. Incorporating labels can significantly improve the quality of embedding transfer, as labels provide valuable contextual information about the data.
- Increased Detectability: Minimizing only the embedding distribution can lead to gradient manipulations that are easier for detection methods (like SweetGuard or Gradient Screener) to identify, as they may not align with the expected gradient behavior for the given labels.
The DAC addresses these limitations by minimizing the joint distribution of embeddings and labels. It is structured to:
- Differentiate Embeddings: Like a standard discriminator, DAC helps differentiate between embeddings generated by the adversary's encoder and those generated by the target model. This drives the target model to produce embeddings that are increasingly similar to the adversary's desired distribution.
- Incorporate Label Information: Crucially, DAC includes an auxiliary classifier that leverages label information. By guiding the target model to not only match the embedding distribution but also to do so in a label-consistent manner, DAC achieves a more effective and nuanced embedding transfer. This ensures that the embeddings produced by the target model not only look similar but also carry the same semantic information as the adversary's encoder.
- Undetectable Gradients: The key benefit of minimizing the joint distribution is that it allows the adversary to generate malicious gradients that are far less detectable. By maintaining consistency with both embedding features and associated labels, the gradients generated by the adversary appear more "honest" or protocol-compliant, effectively bypassing existing gradient-based detection mechanisms.
The speaker presented empirical evidence supporting DAC's superiority. Visualizations showed that the embedding distribution generated from DAC was significantly more similar to the original embedding distribution compared to that from a standard discriminator. Quantitatively, the cosine distance between embeddings was lower with DAC (e.g., 0.02 for MNIST with DAC vs. 0.04 with a standard discriminator), further confirming the improved embedding transfer performance due to label incorporation. This step effectively "poisons" the target model's output embeddings, making them suitable for reconstruction by the adversary's pre-trained decoder.
3. Data Reconstruction
Once the malicious gradient generation phase successfully transfers the embedding distribution, the final step is straightforward: data reconstruction. The adversary uses the pre-trained decoder from the first phase to reconstruct the original private data features from the embeddings received from the passive client. Because the embedding distribution of the target model has been guided to closely mimic that of the adversary's encoder (thanks to DAC), the pre-trained decoder can be highly effective in recovering the data.
The reconstructed data is expected to be of high quality, as the entire URVFL pipeline is optimized for this outcome, from the initial encoder-decoder training to the sophisticated, undetectable embedding transfer mechanism. This technical architecture allows URVFL to achieve its dual goals of high-fidelity reconstruction and stealthy operation, making it a formidable threat in VFL environments.
Demo / Proof of Concept
▶ Watch: Experimental results: URVFL's superior reconstruction performance (8:00)
While the talk did not feature a live, interactive demonstration, the speaker presented compelling empirical validation and visual evidence that served as the proof of concept for URVFL's effectiveness. The "demonstration" was primarily through the rigorous presentation of quantitative results and comparative visualizations across various datasets and attack scenarios.
The empirical validation focused on two key aspects:
- Embedding Transfer Quality: To show how well URVFL's DAC mechanism aligns the target model's embeddings with the adversary's encoder, metrics such as Embedding Mean Squared Error (MSE) distance and Cosine Distance were used. The results clearly indicated that DAC significantly improved embedding transfer, leading to more similar embedding distributions compared to standard discriminator approaches. For instance, a cosine distance of 0.02 was achieved on MNIST with DAC, half of the 0.04 seen with a standard discriminator, directly demonstrating DAC's effectiveness in aligning embedding spaces.
- Data Reconstruction Quality: The ultimate measure of success for URVFL, data reconstruction quality, was evaluated using MSE across all datasets. For image datasets, additional perceptual metrics like Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) were employed to better assess the visual quality of the reconstructed images.
- Tabular Data: On tabular datasets, URVFL consistently outperformed all other baselines. For example, it achieved a reconstruction error of 0.42, which was superior to the suboptimal baseline FSG's 0.5.
- Image Data: For image datasets like CIFAR-10, URVFL (specifically its variant, URVFL-Sync) demonstrated a reconstruction error of 0.01, significantly better than FSG's 0.03. The PSNR and SSIM results also corroborated URVFL's superior reconstruction quality.
Crucially, the "proof of concept" extended to demonstrating URVFL's stealth capabilities under detection. The speaker presented results showing that URVFL maintained its consistent performance even when SweetGuard (SG) and Gradient Screener (GS), two prominent detection methods, were active. In contrast, other malicious attacks like AJN and FSG suffered significant performance degradation under detection, with PSNR values dropping dramatically (e.g., from 22 to 10 for AJN/FSG when facing SG). Furthermore, detection scores were presented, showing that URVFL's score consistently remained above the detection threshold, indicating successful evasion, while other methods fell below, confirming their detectability.
The most compelling part of the demonstration was the visualization of reconstructed images. The speaker displayed reconstructed images from URVFL alongside those from other methods, both with and without active detection. The visualizations clearly showed that URVFL could reconstruct "meaningful pixels" and recognizable images, even when detection mechanisms were engaged. In stark contrast, under detection, other methods like AJN and FSG failed to reconstruct any discernible features, producing largely unintelligible noise. This visual evidence powerfully underscored URVFL's dual capabilities: high-quality reconstruction and robust undetectability.
Defensive Implications
▶ Watch: Crucial finding: URVFL maintains performance under detection (9:20)
The findings presented in the URVFL talk carry profound defensive implications for the design and deployment of Vertical Federated Learning systems. The demonstrated ability of a malicious active client to perform high-fidelity data reconstruction while completely evading current state-of-the-art detection mechanisms highlights a significant blind spot in VFL security.
Firstly, the most immediate implication is that existing gradient-based detection methods, such as SweetGuard and Gradient Screener, are insufficient against sophisticated, malicious adversaries like URVFL. These methods, which rely on identifying anomalies in gradients, are effectively bypassed by URVFL's Discriminator with Auxiliary Classifier (DAC), which crafts gradients that appear benign. Defenders must recognize that simply monitoring gradients for deviations is no longer a sufficient strategy for protecting private data in VFL against determined attackers.
Secondly, this research necessitates a paradigm shift in VFL defense research. Future defensive strategies must move beyond detecting simple gradient anomalies and explore more robust, perhaps multi-faceted, approaches. This could involve:
- Embedding-level scrutiny: Instead of just gradients, more sophisticated analysis of the embeddings themselves, perhaps looking for subtle statistical patterns or information leakage indicators that are independent of gradient structure.
- Protocol-level enhancements: Modifying VFL protocols to introduce cryptographic primitives or zero-knowledge proofs that can verify the honesty of computations without revealing sensitive information.
- Adversarial training for defense: Training VFL systems with defenses specifically designed to counteract attacks that aim for undetectability, perhaps by introducing noise or obfuscation in a way that is robust to URVFL's techniques.
- Trust frameworks and reputation systems: Developing mechanisms to assess and manage the trustworthiness of active clients, although this often moves beyond purely technical solutions.
Thirdly, VFL practitioners and developers must re-evaluate their threat models. The assumption that malicious adversaries are easily detectable by current methods is challenged by URVFL. It is crucial to consider the capabilities of a malicious active client with access to auxiliary data and the intent to stealthily reconstruct private information. This updated threat model should inform the design choices for VFL frameworks, prioritizing defenses against such advanced attacks.
Finally, the talk implicitly calls for proactive security measures. Given the difficulty of detecting URVFL post-factum, preventative measures become even more critical. This might include:
- Data sanitization or differential privacy mechanisms applied at the passive client's end before embeddings are generated, even if they slightly impact model utility.
- Secure aggregation techniques that add layers of privacy protection to the aggregated embeddings or gradients, making it harder for the active client to isolate individual contributions.
In summary, URVFL serves as a stark warning: the current state of VFL security against malicious active clients is precarious. Defenders must urgently innovate and deploy new, more resilient strategies to safeguard the privacy guarantees that VFL purports to offer.
Key Takeaways
- VFL is vulnerable to undetectable data reconstruction attacks: The URVFL attack demonstrates that a malicious active client can reconstruct private data features from passive clients with high fidelity, even when state-of-the-art detection mechanisms are in place.
- Existing gradient-based detection methods are insufficient: Current defenses like SweetGuard and Gradient Screener are bypassed by URVFL, which generates malicious gradients that are indistinguishable from benign ones.
- The Discriminator with Auxiliary Classifier (DAC) is a key innovation: DAC enables stealthy embedding transfer by minimizing the joint distribution of embeddings and labels, which improves transfer quality and crucially reduces detectability compared to standard discriminators.
- URVFL achieves superior reconstruction quality: The attack consistently outperforms baselines (e.g., FSG, AJN) on both tabular and image datasets, achieving lower reconstruction errors (e.g., 0.42 vs 0.5 on tabular, 0.01 vs 0.03 on CIFAR-10) and higher visual quality (PSNR, SSIM).
- A new focus on stealthy malicious adversaries is needed for VFL defense: The research highlights an urgent need for VFL security researchers and practitioners to develop more robust, multi-faceted defensive strategies that account for sophisticated, undetectable attacks from malicious active clients.
- Future defensive work should explore beyond gradient anomaly detection: New defense mechanisms might need to focus on embedding-level scrutiny, cryptographic protocol enhancements, or adversarial training to effectively counteract attacks like URVFL.
About the Speaker(s)
The research on URVFL was presented by Duanyi Yao. While the talk does not provide specific titles or affiliations for Duanyi Yao, it is clear that this work is a significant contribution to the field of federated learning security. Duanyi Yao also acknowledged the joint efforts of Suni Shangongo and Ging Pan, indicating a collaborative research endeavor that led to the development of URVFL. Their work collectively sheds light on critical vulnerabilities within Vertical Federated Learning systems and proposes a novel, sophisticated attack vector that challenges current security assumptions.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid, novel attack research on a real and underexplored threat surface. URVFL's core contribution — using a Discriminator with Auxiliary Classifier to minimize joint embedding-label distribution and generate gradients that evade current detectors — is technically genuine and not something I've seen done cleanly before in the VFL attack literature. The numbers back it up and the threat model is honest about its assumptions.
Heather Calloway (CISO) — WEAK
Technically credible research that demonstrates a real privacy vulnerability in Vertical Federated Learning — but it never escapes the lab. The talk identifies a genuine gap in current VFL defenses and proves it with rigor, but stops well short of telling anyone in a position of authority what to do about it.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025