Please Tell Me More: Privacy Impact of Explainability through the Lens of Membership Inference Attack
Han Liu, Yuhao Wu, Zhiyuan Yu, Ning Zhang
IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 6

Key moments
- 0:00 Introduction: Explainability's challenge in complex AI
- 2:00 Research Question: Does explainability leak privacy?
- 3:10 Key Insight: Attribution maps reveal membership status
- 4:20 Our proposed attack framework for membership inference
- 6:00 Experimental Results: Significant privacy leakage demonstrated
- 7:00 Analyzing the causes of privacy vulnerability
- 7:40 Attack performance under practical, real-world settings
- 8:10 Conclusion and summary of key findings
Please Tell Me More: Privacy Impact of Explainability through the Lens of Membership Inference Attack
Speakers: Han Liu, Yuhao Wu, Zhiyuan Yu, Ning Zhang
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=1mR-TFweLyA
Overview
This research, presented by Han Liu from Washington University in St. Louis, delves into a critical and often overlooked aspect of modern artificial intelligence: the unintended connection between explainable AI (XAI) and privacy leakage. As machine learning models, particularly complex neural networks, become ubiquitous across sensitive domains like healthcare and finance, the demand for transparency and interpretability has soared. XAI methods are designed to demystify these black-box models, providing human-understandable insights into their decision-making processes. However, this talk posits that the very mechanisms intended to provide clarity can inadvertently expose sensitive information about the data used to train these models.
The central thesis of the paper is to systematically investigate whether XAI exacerbates privacy risks, specifically through the lens of membership inference attacks (MIA). MIA aims to determine if a particular data point was part of a model's training dataset. The speakers demonstrate that the additional information provided by XAI, such as attribution maps and prediction trajectories, can be leveraged by adversaries to significantly enhance the efficacy of membership inference attacks. This work highlights a fundamental tension between transparency and privacy in AI, urging a re-evaluation of how XAI systems are designed and deployed, especially in contexts where data privacy is paramount.
The findings presented are significant because they expose a new attack vector that exploits the interpretability features themselves. Given the increasing regulatory and ethical emphasis on both explainability and privacy in AI, understanding this interplay is crucial for developing secure and trustworthy machine learning systems. This research serves as a stark warning and a call to action for researchers and practitioners to consider the privacy implications inherent in the pursuit of model transparency.
Background
▶ Watch: Introduction: Explainability's challenge in complex AI (0:00)
The rapid advancement and widespread deployment of sophisticated neural networks have revolutionized numerous industries and applications. However, as these models grow in complexity and scale, their internal mechanisms become increasingly opaque, leading to a "black-box" problem. This lack of transparency can hinder their adoption, particularly in safety-critical domains like healthcare, autonomous driving, and financial decision-making, where understanding why a model made a certain prediction is as important as the prediction itself. Consequently, explainable AI (XAI) has emerged as a vital field, aiming to provide post-hoc explanations for neural networks, making their operations more understandable to humans.
XAI methods vary widely in their approach and underlying mechanisms. The speakers categorized and focused on four popular types:
- Backpropagation-based methods: These techniques leverage the gradients of model predictions with respect to input features to derive attribution maps, which highlight the most influential parts of an input for a given prediction. Examples include SmoothGrad and ReluGrad.
- Representation-guided methods: These methods employ feature maps from intermediate layers of neural networks to produce attribution maps, often focusing on which parts of the input activate specific internal representations. Grad-CAM and Grad-CAM++ fall into this category.
- Perturbation-based methods: These approaches measure the contribution of each feature by observing the changes in the prediction score when specific features are perturbed or removed. Integrated Gradients (IG) and LIME (Local Interpretable Model-agnostic Explanations) are prominent examples.
- Approximation-based methods: These methods derive explanations by approximating the predictions of the original, complex model with a simpler, inherently interpretable model. SHAP (SHapley Additive exPlanations) is a well-known method in this class.
While XAI addresses a crucial need for transparency, its interaction with data privacy has largely remained underexplored. This talk bridges that gap by investigating XAI's impact on membership inference attacks (MIA). A membership inference attack is a privacy breach where an adversary attempts to determine if a specific data point was included in the training dataset of a machine learning model. In a standard MIA setting, the adversary typically has black-box access to the target model, meaning they can query it with input samples and observe its prediction results (e.g., confidence scores, predicted classes). Membership is often inferred by computing the prediction loss for a given sample: if the loss is unusually small, the sample is likely a member of the training set; otherwise, it is considered a non-member. This vulnerability arises because models tend to generalize better to unseen data but memorize training data, exhibiting a lower loss on samples they have specifically "seen" during training. The core research question addressed in this paper is whether the additional information provided by XAI, beyond standard prediction outputs, can be exploited by an adversary to significantly enhance the success rate of such membership inference attacks, thereby exacerbating privacy leakage.
Key Findings
▶ Watch: Key Insight: Attribution maps reveal membership status (3:10)
The research uncovers a fundamental and exploitable difference in how XAI methods generate explanations for data points that were part of a model's training set (members) versus those that were not (non-members). This disparity forms the cornerstone of their successful membership inference attack.
The central insight is that attribution maps for members tend to emphasize the key semantic features essential for classification more distinctly and consistently compared to those for non-members. For instance, when classifying an image of a cat, the attribution map for a training sample (member) might strongly highlight the cat's eyes, ears, and whiskers, which are definitive features. In contrast, for a non-member image, even if correctly classified, the attribution might be more diffuse, less focused, or highlight less semantically critical areas, indicating a less "learned" or "memorized" feature importance. The speakers visually demonstrated this difference with figures showing original images, followed by attribution maps for members and non-members, illustrating clear visual distinctions in feature focus.
Building upon this insight, the adversary can strategically perturb images guided by these explanation models. By observing how the model's predictions (specifically, confidence scores) change in response to these targeted perturbations, distinct prediction trajectories emerge. Members, having been "memorized" by the model, exhibit a different sensitivity to the perturbation of their key semantic features compared to non-members. This leads to a clear confidence score gap between the perturbation trajectories of members and non-members, which is a critical factor driving the success of the proposed attack.
The comprehensive experimental evaluation yielded several significant findings:
- Significant Privacy Leakage: The proposed attack, leveraging XAI explanations, consistently achieves a better Area Under the Curve (AUC) and a higher True Positive Rate (TPR) at a low False Positive Rate (FPR) (e.g., 0.1% FPR) compared to baseline MIA methods that do not use explanations, as well as other MIA methods that attempt to use explanations but less effectively. This quantitatively demonstrates that XAI methods indeed introduce significant privacy vulnerabilities.
- Broad Applicability: The attack framework is effective across a wide variety of popular explanation methods, including SmoothGrad, ReluGrad, Grad-CAM, IG, Grad-CAM++, 'n' (as listed in the transcript), and SHAP. It also works across different benchmark datasets such as CIFAR-10, CIFAR-100, SVHN, and GTSRB, and various model architectures like ResNet-18 and ResNet-56. This indicates the vulnerability is not specific to a niche XAI method or model type but is a more general phenomenon.
- Root Cause Analysis: The core reason for this privacy vulnerability is attributed to the non-robustness of member samples when their key features are removed or perturbed. This non-robustness, coupled with the distinct perturbation trajectories for each class, exposes different levels of privacy leakage.
- Practicality of Attack: The research also assessed the attack's performance under more practical, real-world conditions. They found that a distribution shift between the shadow dataset (used by the adversary to train their attack model) and the target model's training data had only a minor influence on the attack's performance. Furthermore, architectural differences between the adversary's shadow model and the target model also had a limited impact on attack effectiveness. These findings suggest that the attack is robust and can be effective even when the adversary does not have perfect knowledge or resources aligning with the target.
In summary, the key findings unequivocally demonstrate that XAI, while beneficial for transparency, introduces a significant and practical privacy risk by providing adversaries with additional, exploitable signals that enhance membership inference attacks.
Technical Deep Dive
▶ Watch: Experimental Results: Significant privacy leakage demonstrated (6:00)
The adversary's goal in this enhanced membership inference attack is to determine if a specific data point was used to train a target machine learning model. Unlike standard MIA, this framework assumes the adversary has black-box access to the target model's explanations in addition to its prediction results. This crucial assumption allows the adversary to leverage the insights provided by XAI methods.
The attack framework is structured in four distinct stages:
- Stage 1: Shadow Model Training
The adversary first trains a shadow model. This shadow model is designed to mimic the behavior of the target model. The adversary typically trains this model on a proxy dataset drawn from the same general distribution as the target model's training data, though the research shows this doesn't need to be an exact match. The purpose of the shadow model is to simulate the target model's responses, including its explanations, which the adversary needs to generate the training data for their own attack model.
- Stage 2: Explanation Generation and Perturbation Trajectory Measurement
In this stage, the adversary queries the trained shadow model. For each sample (both those known to be "members" of the shadow model's training set and "non-members"), the adversary utilizes an explanator (an XAI method) to generate attribution maps. These maps highlight the most salient features of the input image that contributed to the model's prediction.
Crucially, the adversary then applies different levels of perturbation to the sample, guided by these attribution maps. This means features identified as highly important by the attribution map might be systematically altered, removed, or masked. For each perturbation level, the adversary measures the changes in the model's prediction (specifically, the confidence score for the predicted class). The sequence of confidence scores observed across these increasing perturbation levels forms a prediction trajectory.
The core insight here is that the prediction trajectories for member samples, which the model has "memorized," will differ observably from those of non-member samples. For instance, removing a key feature from a member sample might cause a sharp drop in confidence, whereas the same perturbation on a non-member might have a less dramatic effect or follow a different pattern. The speakers highlight that hypothesis testing is used to select the most informative trajectories, ensuring that the collected data best distinguishes between members and non-members.
- Stage 3: Feature Aggregation and Membership Feature Construction
Once the prediction trajectories are generated, the adversary extracts and aggregates various features to construct a comprehensive set of membership features for each sample. These features include:
- Attribution features: Derived from the attribution maps themselves, capturing the patterns of feature importance.
- Prediction trajectory features: Quantifying the observed changes in confidence scores under perturbation.
- Loss: The standard prediction loss, which is a common feature in traditional MIA.
- One-hot encoding of the predicted classes: Providing contextual information about the model's output.
These aggregated features form a rich representation for each sample, designed to capture the subtle differences in model behavior between members and non-members when XAI is involved.
- Stage 4: Attack Model Training and Membership Inference
In the final stage, the adversary trains a binary classifier (referred to as the "attack model," e.g., an ordinary perceptron or a small neural network) using the aggregated membership features derived from the shadow model. The attack model learns to differentiate between members and non-members based on these features.
Finally, to infer the membership status of a target sample, the adversary inputs the corresponding membership features (generated by querying the target model and its explanator) into the trained attack model. The attack model then outputs a prediction indicating whether the target sample was a member or a non-member of the target model's training dataset.
The research conducted extensive evaluations to validate this framework. They tested seven popular explanation methods: SmoothGrad, ReluGrad, Grad-CAM, IG, Grad-CAM++, 'n' (as stated in the transcript), and SHAP. The attack was evaluated across four benchmark datasets: CIFAR-10, CIFAR-100, SVHN, and GTSRB, and against different model architectures, specifically ResNet-18 and ResNet-56. The performance was measured using standard metrics like Area Under the Curve (AUC) of the ROC curve and the True Positive Rate (TPR) at a False Positive Rate (FPR) of 0.1%. The results consistently showed a significant improvement in attack performance when explanations were leveraged, highlighting the enhanced privacy leakage.
Furthermore, the study delved into the practical implications of the attack. It was found that a distribution shift in the shadow dataset (the data used to train the adversary's attack model) had only a minor impact on performance. Similarly, architectural differences between the adversary's shadow model and the actual target model did not significantly degrade the attack's effectiveness. These findings underscore the robustness and real-world applicability of this XAI-enhanced membership inference attack. The core technical takeaway is that the unique "fingerprint" left by member samples in their attribution maps and perturbation trajectories, particularly their non-robustness to removal of key features, is a powerful signal for membership inference.
Demo / Proof of Concept
▶ Watch: Analyzing the causes of privacy vulnerability (7:00)
While the talk did not feature a live, interactive demonstration of the attack framework with running code, the speakers presented compelling visual evidence and experimental results that served as a strong proof of concept for their core hypothesis. Specifically, they utilized illustrative figures to highlight the key observations that underpin their attack.
These figures visually contrasted the attribution maps generated for member samples versus non-member samples. They clearly showed how the attribution maps for members focused more intensely and accurately on the semantically relevant features of an image (e.g., the distinguishing parts of an animal for a classification task), whereas non-member attribution maps were often more diffuse or focused on less critical areas. This visual distinction directly supported their central insight that XAI explanations provide a unique signal for membership.
Furthermore, the presentation included plots illustrating the confidence score gaps and perturbation trajectories for members and non-members. These graphs demonstrated how the model's prediction confidence changed distinctly for members when their crucial features (identified by attribution maps) were progressively perturbed, compared to the changes observed for non-members. This visual evidence, combined with the comprehensive quantitative evaluation metrics (AUC, [email protected]%FPR) presented across various XAI methods, datasets, and model architectures, effectively served as the empirical proof of concept for the feasibility and effectiveness of their XAI-enhanced membership inference attack. The extensive experimental results, rather than a live code demo, formed the primary demonstration of the attack's capabilities.
Defensive Implications
▶ Watch: Conclusion and summary of key findings (8:10)
The findings of this research carry significant implications for developers, deployers, and users of machine learning models, particularly those incorporating XAI. The revelation that XAI can inadvertently exacerbate privacy leakage necessitates a proactive and multi-faceted approach to defense.
Firstly, model developers and MLOps teams must recognize that the pursuit of interpretability can introduce new privacy vulnerabilities. When designing and implementing XAI features, privacy considerations should be integrated from the outset, rather than being an afterthought. This calls for the development of privacy-preserving XAI methods that can provide explanations without revealing sensitive training data information. This might involve techniques that obfuscate or generalize explanations for individual data points, or that are inherently designed with differential privacy guarantees.
Secondly, organizations deploying AI models in sensitive domains (e.g., healthcare, finance, legal systems) must exercise extreme caution when also exposing XAI features. If XAI outputs are accessible to external parties, even in a black-box query setting, they become potential vectors for privacy attacks. Risk assessments for AI systems should explicitly include an evaluation of XAI features for membership inference risks. It might be necessary to restrict access to raw explanation outputs or provide only aggregated, anonymized explanations to end-users.
Thirdly, defensive strategies commonly used against traditional membership inference attacks could be adapted or enhanced. Differential privacy (DP) is a strong candidate for mitigating this risk. By introducing carefully calibrated noise during model training or during the generation of explanations, DP can provide provable privacy guarantees, making it difficult for an adversary to infer membership. However, applying DP to XAI methods while retaining their explanatory power remains an active area of research.
Fourthly, adversarial training techniques could be explored. If models can be trained to be robust against perturbations that reveal membership status through XAI, this could offer a defense. This would involve training the model with a loss function that penalizes not only incorrect predictions but also the leakage of membership information through its explanations. The research notes that member samples are "non-robust" to the removal of key features; a defense could aim to make them more robust, or conversely, make non-members similarly non-robust to equalize the signal.
Finally, there is a clear need for auditing and validation of XAI systems for privacy vulnerabilities. Before deploying an XAI-enabled model, it should be rigorously tested for susceptibility to membership inference attacks that leverage its explanation capabilities. This could involve red-teaming exercises where ethical hackers attempt to infer membership using the XAI outputs, providing valuable feedback for system hardening. The findings also suggest that simply having a shadow model that differs architecturally or is trained on a shifted distribution is not sufficient to prevent the attack, meaning defenses need to be more fundamental.
In essence, the research highlights that transparency and privacy are not always complementary goals. Achieving both requires a deliberate and sophisticated approach, moving beyond simply applying XAI to black-box models and towards developing privacy-aware XAI systems from the ground up.
Key Takeaways
- XAI Exacerbates Privacy Leakage: Explainable AI methods, while intended to increase transparency, inadvertently create new avenues for privacy breaches, specifically by enhancing the efficacy of membership inference attacks.
- Distinct Explanations for Members: Attribution maps generated by XAI methods for training data members are demonstrably more focused on key semantic features essential for classification compared to those for non-members, providing a critical signal for adversaries.
- Exploitable Prediction Trajectories: By perturbing images guided by XAI attribution maps, adversaries can observe unique "prediction trajectories" and "confidence score gaps" between members and non-members, enabling their differentiation.
- Robust and Broadly Applicable Attack: The proposed attack framework is effective across a wide array of popular XAI methods (e.g., SmoothGrad, Grad-CAM, SHAP), benchmark datasets (e.g., CIFAR-10, SVHN), and model architectures (e.g., ResNet-18, ResNet-56).
- Practical Attack Efficacy: The attack's performance remains largely unaffected by practical considerations such as distribution shifts in the adversary's shadow dataset or architectural differences between the shadow and target models, underscoring its real-world applicability.
- Urgent Need for Privacy-Aware XAI: There is a critical and immediate need for the development and deployment of privacy-preserving XAI techniques and for integrating privacy risk assessments into the design and deployment of any AI system that exposes interpretability features.
About the Speaker(s)
The presentation was delivered by Han Liu from Washington University in St. Louis. Han Liu is listed as the primary speaker, sharing the research findings. The paper itself is a collaborative effort, with Yuhao Wu, Zhiyuan Yu, and Ning Zhang also listed as co-authors on the research. The talk highlights the work conducted at Washington University in St. Louis, focusing on the intersection of machine learning, security, and privacy.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This research meticulously demonstrates how Explainable AI (XAI) methods, intended for transparency, inadvertently amplify privacy risks by providing new, exploitable signals for membership inference attacks. The attack leverages distinct attribution maps and prediction trajectories for training members versus non-members, a clever and concerning new vector. It's a critical warning that transparency and privacy are often at odds in ML systems.
Heather Calloway (CISO) — STRONG ACCEPT
This research uncovers a significant and practical privacy risk inherent in Explainable AI, demonstrating how XAI outputs can be exploited to enhance membership inference attacks. It forcefully highlights the critical tension between transparency and privacy, demanding that security leaders reassess their approach to AI governance and deployment.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024