Probe-Me-Not: Protecting Pre-trained Encoders from Malicious Probing
Ruyi Ding (PhD student · Northeastern University)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · ML Security
Overview
In the rapidly evolving landscape of machine learning, the paradigm of transfer learning has become a cornerstone, enabling the development of highly accurate models with significantly reduced data and computational resources. This talk, "Probe-Me-Not: Protecting Pre-trained Encoders from Malicious Probing," delivered by Ruyi Ding from Northeastern University, addresses a critical intellectual property (IP) and security threat inherent in the widespread use of pre-trained models, particularly those offered as API services. The core problem revolves around preventing the misuse of these powerful pre-trained encoders for "prohibited" or "harmful" tasks, even as they continue to perform optimally for their intended, "authorized" applications.
Key moments
- 0:00 Malicious probing of pre-trained encoders problem
- 2:20 Encoder Locker framework for restricted transferability
- 3:25 Three core design objectives of the framework
- 4:40 Domain-aware weight optimization and new loss
- 6:00 Robust protection through adversarial training
- 7:00 Addressing prohibited data accessibility (supervised case)
Probe-Me-Not: Protecting Pre-trained Encoders from Malicious Probing
Speakers: Ruyi Ding (PhD student, Northeastern University)
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=GG2YB1OuzM4
Overview
In the rapidly evolving landscape of machine learning, the paradigm of transfer learning has become a cornerstone, enabling the development of highly accurate models with significantly reduced data and computational resources. This talk, "Probe-Me-Not: Protecting Pre-trained Encoders from Malicious Probing," delivered by Ruyi Ding from Northeastern University, addresses a critical intellectual property (IP) and security threat inherent in the widespread use of pre-trained models, particularly those offered as API services. The core problem revolves around preventing the misuse of these powerful pre-trained encoders for "prohibited" or "harmful" tasks, even as they continue to perform optimally for their intended, "authorized" applications.
The presentation introduces a novel protective framework named Encoder Locker, designed to proactively restrict the transability of pre-trained models to undesirable tasks. This is a crucial area of research known as applicability authorizations, which formally aims to prevent models from being misused by proactively restricting their adaptability for harmful objectives. Ding's work focuses on a flexible attacker model, ensuring that while the encoder provides high accuracy for legitimate downstream tasks, it yields very low accuracy for prohibited tasks, effectively making the embeddings useless for malicious purposes. This dual objective of maintaining authorized performance while degrading prohibited performance forms the bedrock of the Encoder Locker framework, presenting a significant advancement in safeguarding the integrity and intended use of valuable pre-trained assets.
Background
▶ Watch: Malicious probing of pre-trained encoders problem (0:00)
The modern machine learning ecosystem heavily relies on transfer learning, a technique where models are not trained from scratch but rather fine-tuned from existing, powerful pre-trained checkpoints. This approach drastically reduces the data requirements and training costs for users, democratizing access to advanced AI capabilities. A common instantiation of transfer learning involves freezing the initial, complex layers—often convolutional or attention layers—which act as a feature encoder. Users then attach their own simple, linear downstream components, a process commonly referred to as linear probing or model probing, to adapt the encoder's embeddings for specific tasks. Many companies now offer these pre-trained encoders as API services, allowing users to query the API, obtain embeddings, and build highly accurate models on their own datasets with minimal effort.
While immensely beneficial, this accessibility introduces a significant vulnerability: the potential for misuse. The problem arises when a malicious actor attempts to redirect these pre-trained encoders for tasks that are not permitted or are explicitly prohibited by the API provider. For instance, a general-purpose image encoder, designed for benign classification tasks, could potentially be repurposed for identifying military equipment or other sensitive content, violating ethical guidelines or terms of service. Previous research in applicability authorizations has acknowledged this challenge, defining it as the proactive restriction of a model's transability to harmful tasks. However, many prior works struggle with the ambiguity of defining "prohibited data" and ensuring the robustness of protections against sophisticated attackers who might modify their downstream heads. The challenge is to create a mechanism that ensures the pre-trained encoder provides high-quality embeddings for authorized tasks (source domains) but generates embeddings that lead to very low accuracy for unauthorized tasks (target domains), regardless of the attacker's downstream model architecture or training data. This requires a robust, flexible, and adaptable defense mechanism that can operate even when explicit labels or examples of prohibited data are scarce or non-existent.
Key Findings
▶ Watch: Three core design objectives of the framework (3:25)
The "Probe-Me-Not" research introduces Encoder Locker, a protective framework designed to make pre-trained encoders robust against malicious probing for prohibited tasks. The core findings center around its ability to achieve three critical design objectives:
- Interactivity and Restrictions: Encoder Locker ensures that the performance of the pre-trained model on authorized source domains remains high, ideally as good as its original unprotected state, while simultaneously driving down the performance on target domains (prohibited tasks) to a very low accuracy. This effectively restricts the transability of the encoder to harmful applications.
- Robustness Against Flexible Attackers: Recognizing that attackers can adjust their downstream heads and training data in various ways, Encoder Locker incorporates a self-challenge method. This adversarial training approach ensures that the protection remains robust against adaptive attackers, making it difficult for them to circumvent the restrictions by simply changing their downstream model architecture or optimization strategy.
- Prohibited Data Accessibility: Addressing a major gap in previous work, Encoder Locker offers three distinct levels of protection based on the availability of information about the prohibited domains:
- Supervised Encoder Locker: When clear labels and data for prohibited domains are available.
- Unsupervised Encoder Locker: When only data samples from prohibited domains are available, but without explicit labels.
- Zero-Shot Encoder Locker: When neither data nor labels are available, relying solely on textual descriptions of prohibited domains.
The evaluation demonstrates that Encoder Locker consistently achieves a very high accuracy drop on target domains across all three levels of protection, while maintaining strong performance on source domains. The research also provides compelling visualizations of the latent space, showing clear separation for authorized data but a highly mixed, uninformative representation for prohibited data, confirming the framework's effectiveness in degrading the utility of embeddings for malicious use.
Technical Deep Dive
▶ Watch: Domain-aware weight optimization and new loss (4:40)
The Encoder Locker framework is built upon a sophisticated technical architecture designed to achieve its three primary objectives.
Objective 1: Interactivity and Restrictions (Domain-Aware Weight Optimizations)
At the heart of Encoder Locker's ability to selectively degrade performance is its domain-aware weight optimization. The fundamental idea is to identify and modify specific weights within the pre-trained encoder that are critical to the target (prohibited) domain, without negatively impacting performance on the source (authorized) domain.
The optimization process is guided by a novel loss function called encoder_locker_loss, defined as:
encoder_locker_loss = original_loss_source - log(target_loss)
- The
original_loss_sourceterm aims to keep the loss on the source domains very low, thereby ensuring that the encoder maintains high accuracy for its intended tasks. - The
log(target_loss)term serves as a regularization component. By maximizing the logarithm of the target loss (equivalent to minimizing-log(target_loss)), the framework actively works to make the target loss as high as possible. The use of the logarithm ensures that this regularization term remains effective even when the target loss becomes very high, preventing numerical instability or saturation.
The speaker highlights that the challenge lies in identifying "what to optimize" – finding those specific weights that are critical to the target domain's performance. The paper, though not detailed in the talk, specifies a method for identifying these crucial weights, allowing for targeted modification rather than a global re-training of the entire encoder.
Objective 2: Robustness Against Flexible Attackers (Self-Challenge Method)
To ensure the protection is robust against a flexible attacker who can adapt their downstream head, Encoder Locker employs a self-challenge method, conceptually similar to adversarial training. This involves an iterative, two-step process:
- Attacker Adaptation: In one step, the framework simulates an attacker by training a downstream head that attempts to maximize the
encoder_locker_lossbased on the current state of the protected encoder. This represents the attacker trying to find a way to extract useful information from the encoder for the prohibited task. - Encoder Adaptation: In the subsequent step, the encoder itself is updated. It selects a new set of weights and optimizes them to minimize the
encoder_locker_loss, effectively counteracting the attacker's adaptation.
This iterative adversarial process forces the encoder to learn robust protections that are resilient to various downstream head configurations an attacker might employ, thus preventing easy circumvention of the security measures.
Objective 3: Prohibited Data Accessibility (Three Levels of Encoder Locker)
Addressing the practical challenge of defining and accessing prohibited data, Encoder Locker offers three flexible levels of protection:
- Supervised Encoder Locker:
- Scenario: This level applies when a clear dataset with explicit labels for the prohibited (target) domains is available.
- Mechanism: The framework directly incorporates a standard cross-entropy loss (or a similar classification loss) for the target domain into the
encoder_locker_lossfunction. This allows the system to explicitly learn to misclassify or degrade features for known prohibited categories.
- Unsupervised Encoder Locker:
- Scenario: This level is designed for situations where data samples from prohibited domains are available, but they lack explicit labels. For example, a provider might collect images known to be problematic but not have the resources to manually label them.
- Mechanism: It leverages principles from contrastive learning. For a given batch of prohibited data, it constructs:
- Positive pairs: Augmentations of the same data sample.
- Negative pairs: Different data samples within the batch.
- The objective is to minimize the similarity between negative pairs and maximize the similarity between positive pairs within the context of the target loss. When plugged into the
encoder_locker_loss, this essentially forces the encoder to produce embeddings for prohibited data that are poorly clustered or indistinguishable, making it difficult for any downstream head to differentiate or classify them meaningfully.
- Zero-Shot Encoder Locker:
- Scenario: This is the most challenging scenario, where neither data samples nor labels for prohibited domains are available. The only information is a textual description of the prohibited domain (e.g., "military usage").
- Mechanism: This level innovatively leverages generative models and AI agents.
- An AI agent uses the textual description to generate a synthetic data set for the prohibited domain, often employing text-to-image generators.
- A "refine prompt algorithm" is then used to refine the synthetic data set, ensuring it adequately covers the described prohibited domains.
- Once the synthetic data is generated, it is treated similarly to the unsupervised scenario, applying a contrastive loss or similar technique to ensure the generated data's embeddings are degraded. This allows protection against entirely novel or abstractly defined prohibited domains.
In summary, the complete Encoder Locker framework integrates a domain-aware weight search algorithm to identify critical weights, a self-challenge method for robustness, and these three flexible levels of protection, making it a comprehensive solution for safeguarding pre-trained encoders against diverse forms of malicious probing.
Demo / Proof of Concept
▶ Watch: Robust protection through adversarial training (6:00)
The presentation provided compelling evidence of Encoder Locker's effectiveness through several demonstrations and visualizations:
- Accuracy Drop Curves: The core quantitative evaluation involved plotting accuracy drop curves for both source (authorized) and target (prohibited) domains.
- The "right curve" represented the performance on source domains, showing the degradation (or lack thereof) when Encoder Locker was applied. Critically, this curve demonstrated minimal accuracy drop, indicating that the protection mechanism does not significantly harm the model's performance on its intended tasks.
- The "blue curve" illustrated the accuracy drop on target domains. For all three levels—supervised, unsupervised, and zero-shot—a very high accuracy drop was observed. This signifies that the embeddings generated for prohibited data become largely useless for downstream classification, fulfilling the primary objective of restricting transability.
- It was noted that the zero-shot level, while effective, exhibited a slightly less pronounced accuracy drop compared to supervised and unsupervised methods. This is attributed to the inherent challenges in generating perfectly representative synthetic data via AI agents, which cannot always guarantee that the generated data perfectly covers the target domain without some unintended impact.
- Latent Space Visualization: To provide a deeper understanding of why the accuracy drops, the talk presented visualizations of the latent space (the embedding space) for both source and target domains.
- For source domains, the embeddings remained well-separated, indicating that the encoder still produces distinct and discriminative features for authorized categories. A downstream head can easily classify these.
- In stark contrast, the embeddings for target domains were mixed up together. This "mixing" means that the encoder no longer provides useful, separable information for prohibited classes. Any downstream classification head would struggle to differentiate between these mixed-up embeddings, leading to the observed low accuracy.
- Military Data Set Example: A concrete example was provided using a "military data set" as a prohibited domain, featuring images of tanks.
- For an unprotected encoder, the visualization highlighted the "tank gun" as the most important feature (red highlight) for classification, indicating the model's ability to accurately identify military equipment.
- When supervised Encoder Locker was applied, the critical features shifted to the "wheels" of the tank. This misdirection causes the model to misclassify, for example, a tank as a combat vehicle or some other wheeled object, effectively making the original classification task fail.
- For unsupervised and zero-shot Encoder Locker, the feature highlights became "out of focus" or dispersed across the image. This indicated that the encoder could no longer extract any clear, useful, or discriminative information from the prohibited images, rendering them effectively unclassifiable for the intended malicious purpose.
These demonstrations collectively provided strong empirical evidence that Encoder Locker successfully restricts the transability of pre-trained encoders to prohibited tasks while preserving performance on authorized applications, even under various levels of information availability regarding the prohibited domains.
Defensive Implications
▶ Watch: Addressing prohibited data accessibility (supervised case) (7:00)
The "Probe-Me-Not" research offers profound defensive implications for organizations and API providers offering pre-trained machine learning models. In an era where powerful foundation models are increasingly accessible as services, ensuring their ethical and intended use is paramount. Encoder Locker provides a concrete, multi-faceted framework to achieve this.
Firstly, model providers can proactively embed these protections into their pre-trained encoders before deployment. This shifts the burden of misuse prevention from reactive measures (e.g., monitoring API queries for suspicious patterns) to a proactive, architectural defense. By making the embeddings for prohibited tasks effectively useless, providers can mitigate the risk of intellectual property theft, ethical violations, or regulatory non-compliance.
Secondly, the framework's flexibility, particularly its three levels of protection (supervised, unsupervised, and zero-shot), is a significant advantage. Organizations may not always have perfectly labeled datasets of "bad" behavior.
- For well-defined prohibited categories with available data and labels (e.g., hate speech, specific illegal content), the supervised Encoder Locker offers direct and robust protection.
- When only examples of prohibited content are available without explicit labels (e.g., a collection of sensitive images), the unsupervised Encoder Locker leverages contrastive learning to make these embeddings indistinguishable.
- Most powerfully, the zero-shot Encoder Locker allows providers to define prohibited domains purely through textual descriptions (e.g., "models must not be used for military applications," "models must not be used for surveillance"). This capability to generate synthetic prohibited data and enforce restrictions without any prior examples is revolutionary for addressing emerging threats or vaguely defined misuse cases.
A key consideration, highlighted during the Q&A session, is the scenario of multiple target domains. While the framework is adaptable, simultaneously protecting against numerous distinct prohibited tasks might lead to a slight degradation in performance on the source domain. Defenders need to carefully weigh the trade-offs between the number and specificity of prohibited tasks and the desired performance on authorized tasks. This suggests a strategic approach where the most critical or high-risk prohibited domains are prioritized for protection.
Ultimately, Encoder Locker empowers model providers to enforce "applicability authorizations" directly within the model's architecture, ensuring that their valuable pre-trained assets are used responsibly and in alignment with their intended purpose, thereby safeguarding both their IP and their ethical standing.
Key Takeaways
- Addressing IP Threat: Pre-trained encoders, widely used via transfer learning and linear probing, face an IP threat where malicious actors can repurpose them for prohibited tasks.
- Encoder Locker Framework: This novel framework proactively restricts the transability of pre-trained models to harmful tasks while maintaining high performance on authorized tasks.
- Three Core Objectives: The framework achieves interactivity and restriction, robustness against flexible attackers via a self-challenge method, and adaptable prohibited data accessibility.
- Flexible Protection Levels: Encoder Locker offers supervised (with labels), unsupervised (data only), and zero-shot (textual description only, using generative AI) protection levels to accommodate varying data availability for prohibited domains.
- Demonstrated Effectiveness: Evaluations show a significant accuracy drop on target (prohibited) domains with minimal impact on source (authorized) domains, confirmed by latent space visualizations and specific examples like the military dataset.
- Proactive Defense: It provides a powerful tool for model providers to embed ethical and usage restrictions directly into their models, mitigating misuse and protecting intellectual property.
About the Speaker(s)
Ruyi Ding is a PhD student at Northeastern University. His research, presented in "Probe-Me-Not," focuses on critical areas of machine learning security, specifically protecting pre-trained models from malicious probing and misuse. He collaborated with a team of professors from Northeastern University, including Professor Lisu, Professor Adam, Professor, and Professor Yuf, and acknowledged the support of the National Science Foundation for their project.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic ML security research with a coherent threat model and a technically sound framework. The problem — preventing malicious reuse of pre-trained encoders via linear probing — is real and undersolved. But this is a venue paper presentation, not a practitioner security talk, and the gap between academic elegance and deployment reality is never closed.
Heather Calloway (CISO) — WEAK
Technically coherent research on protecting pre-trained encoders from misuse, but it never makes the institutional leap. The work may be sound; the talk is built for a ML security audience, not the operators or leaders who actually govern AI deployment risk.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025