GrOVe: Ownership Verification of Graph Neural Networks using Embeddings
Asim Waheed, Vasisht Duddu, N. Asokan
IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5
Overview
In an era where Graph Neural Networks (GNNs) are becoming indispensable for modeling complex real-world relationships in social networks, recommendation systems, and scientific applications, the intellectual property of these sophisticated models faces significant threats. This talk introduces GrOVe (Ownership Verification of Graph Neural Networks using Embeddings), a novel framework designed to demonstrate and verify the ownership of GNNs in the face of increasingly sophisticated model extraction attacks. Presented by Asim Waheed, alongside co-authors Vasisht Duddu and N. Asokan, the research addresses a critical vulnerability in the deployment of proprietary GNN models.

Key moments
- 0:00 Introduction: GNNs and ownership verification need
- 1:30 Understanding GNN Model Extraction Attacks
- 2:50 Key insight: GNN embeddings as model fingerprints
- 3:30 Desired properties of ownership verification scheme
- 4:20 Threat model and four key stakeholders defined
- 5:50 Model registration with cryptographic commitments
- 7:20 Proposed verification algorithm using similarity classifier
GrOVe: Ownership Verification of Graph Neural Networks using Embeddings
Speakers: Asim Waheed; Vasisht Duddu; N. Asokan
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=If1s7We2gAI
Overview
In an era where Graph Neural Networks (GNNs) are becoming indispensable for modeling complex real-world relationships in social networks, recommendation systems, and scientific applications, the intellectual property of these sophisticated models faces significant threats. This talk introduces GrOVe (Ownership Verification of Graph Neural Networks using Embeddings), a novel framework designed to demonstrate and verify the ownership of GNNs in the face of increasingly sophisticated model extraction attacks. Presented by Asim Waheed, alongside co-authors Vasisht Duddu and N. Asokan, the research addresses a critical vulnerability in the deployment of proprietary GNN models.
The core problem GrOVe tackles is the ease with which an adversary can extract a functional replica of a deployed GNN by interacting with its prediction API, thereby circumventing the extensive resources and expertise invested in its development. The authors demonstrate that such "surrogate" models, despite being independently trained, generate remarkably similar node embeddings to the original "target" models. GrOVe leverages this fundamental insight, proposing a robust and efficient method for a third-party verifier to ascertain whether a suspect GNN has been illicitly extracted from a legitimate target model, thereby safeguarding the intellectual property of GNN developers and deployers.
Background
▶ Watch: Introduction: GNNs and ownership verification need (0:00)
Graph Neural Networks (GNNs) represent the state-of-the-art for applications built upon graph-structured data, enabling tasks like node classification, graph classification, visualization, and recommendations. The fundamental principle of a GNN involves aggregating information from a target node's neighbors, combining node features and graph structure to produce a single, rich embedding vector for each node. These embeddings capture the node's contextual information within the graph, which can then be fed into subsequent layers or other neural networks for various downstream tasks.
The increasing prevalence and commercial value of GNNs have, however, attracted the attention of adversaries seeking to exploit them. A significant threat, as highlighted by prior work, is model extraction. In a typical model extraction attack, an adversary interacts with a target model via a prediction API. By sending carefully crafted inputs and observing the model's outputs, the adversary can train their own "surrogate" model that mimics the behavior of the original target model. This surrogate model often performs comparably to the target but is significantly cheaper for the adversary to obtain, bypassing the original training costs and data acquisition efforts.
Specifically for GNNs, Shen et al. demonstrated the feasibility of model extraction, particularly for inductive GNNs. They identified two primary attack types:
- Type 1 Attack: The adversary has access to both the adjacency matrix (representing graph edges) and the node features, directly training a surrogate GNN.
- Type 2 Attack: The adversary does not have the adjacency matrix and must first estimate it before training the surrogate.
In both scenarios, the adversary's objective is two-fold: achieve high accuracy on the primary task (e.g., node classification) and maintain high fidelity between the target and surrogate models, ensuring similar output behavior.
The core insight underpinning GrOVe stems from this fidelity requirement. If a surrogate model accurately mimics a target GNN, it stands to reason that the embeddings generated by both models for the same input graph should be highly similar. A preliminary experiment using t-SNE visualization confirmed this intuition: embeddings from surrogate models fully overlapped with those from the target model, while embeddings from independently trained models (identical architecture but different random initialization) formed distinct clusters. This visual evidence strongly suggested that GNN embeddings could serve as a unique fingerprint for ownership verification.
Before developing GrOVe, the authors established four crucial desiderata for any effective ownership verification scheme:
- Effective: Must reliably differentiate between surrogate and independently trained models.
- Efficient: Should incur reasonable computational overhead.
- Non-invasive: Must not degrade the target model's accuracy (a common requirement for fingerprinting).
- Robust: Must resist adversarial attempts to circumvent detection, such as model compression or fine-tuning.
The threat model for GrOVe aligns with existing model extraction literature. The adversary is assumed to have black-box access to the target GNN's output, specifically the node embeddings, which are used to train their surrogate. The overlap between the training datasets of the target and surrogate models can range from none to full; GrOVe's experiments specifically focused on the most challenging scenario: full data overlap. For the verifier, a black-box setting is also assumed: the verifier can sample a verification dataset from the same distribution as the target model's training data but does not require access to the exact original data. The verifier can also access both the target and suspect models via prediction APIs.
To formalize the verification process, GrOVe defines four stakeholders:
- Model Owner: Trains and deploys the GNN service.
- Adversarial Responder: Steals a model and attempts to evade detection.
- Adversarial Accuser: May or may not have their own model but falsely accuses an independent model owner of theft.
- Third-Party Verifier: A trusted entity responsible for determining if a model was stolen.
To prevent false accusations and establish a clear timeline, GrOVe introduces a model registration process. When a model owner trains and deploys a GNN, they must generate a cryptographic commitment (C) of their model (e.g., a simple cryptographic hash of its weights or architecture). This commitment must change if the model is altered. Crucially, the owner must then obtain a secure timestamp for C and provide both to the trusted verifier. This registration step ensures the verifier knows which model was trained first, laying the groundwork for verifiable ownership claims.
Key Findings
▶ Watch: Key insight: GNN embeddings as model fingerprints (2:50)
The research presented on GrOVe yielded several significant findings regarding the feasibility and effectiveness of GNN ownership verification:
- Embedding Similarity as a Fingerprint: The foundational discovery is that surrogate GNN models, extracted from a target GNN, produce remarkably similar node embeddings to the original target model. This similarity is robust enough to serve as a reliable fingerprint for ownership verification, as evidenced by initial t-SNE visualizations showing complete overlap between target and surrogate embeddings, distinct from independently trained models.
- High Effectiveness (Low Error Rates): GrOVe demonstrated exceptional effectiveness in identifying stolen GNNs. It achieved a zero percent false negative rate (FNR) across all tested datasets (Amazon, ACM, Co-Author, etc.) for both Type 1 and Type 2 model extraction attacks. This is particularly noteworthy as the similarity classifier was only trained on Type 1 attack data, showcasing its generalization capabilities. Furthermore, GrOVe maintained very low false positive rates (FPR), with only 2.2% for the Amazon dataset and 3.4% for the ACM dataset. For all other datasets, GrOVe achieved an impressive 0% FPR, meaning no independently trained models were falsely accused of being stolen.
- Robustness Against Double Extraction: GrOVe proved robust against sophisticated evasion techniques like double extraction, where an adversary extracts an intermediate model and then extracts a surrogate from that intermediate model. Despite the added layer of knowledge distillation, GrOVe consistently maintained a 0% FNR across all datasets. Interestingly, double extraction often led to a decrease in the surrogate model's accuracy, with drops of over 6% observed in the Amazon and ACM datasets, indicating it's not only less potent but also still effectively caught by GrOVe.
- Addressing Pruning with Data Augmentation: While initial testing showed that pruning (randomly removing model weights) could successfully evade GrOVe when applied with moderate prune ratios (between 0.1 and 0.4), the researchers developed an effective countermeasure. By augmenting the training data of the similarity classifier to include outputs from pruned surrogate models (specifically those with prune ratios up to 0.4), GrOVe regained its robustness. After this augmentation, all pruned surrogate models within this range were correctly identified. Pruning beyond a 0.4 ratio typically resulted in significant accuracy degradation (greater than 10%), rendering the surrogate model less useful to the adversary.
- Computational Efficiency: GrOVe demonstrated practical efficiency. The time required to generate training data for and train the similarity classifier was primarily influenced by the dataset size. For the largest dataset, Co-Author, training completed in less than 3 hours using a single A100 GPU with 40 GB of RAM. This suggests that GrOVe's computational overhead is reasonable for real-world deployment.
Technical Deep Dive
▶ Watch: Desired properties of ownership verification scheme (3:30)
The technical foundation of GrOVe lies in the inherent behavior of GNNs and the specific characteristics of model extraction attacks. A GNN operates by iteratively aggregating information from a node's local neighborhood. Given node features and the graph structure (adjacency matrix), it produces a low-dimensional embedding vector for each node. These embeddings are crucial as they encapsulate the node's structural and semantic context within the graph, making them highly informative for downstream tasks.
Model extraction attacks on GNNs, as identified by Shen et al., exploit the black-box access to a target GNN's prediction API. The adversary queries the target with various inputs and observes the resulting node embeddings. These observed embeddings then serve as "labels" for training the adversary's own surrogate GNN. The success of such an attack hinges on the surrogate's ability to replicate the target's behavior, implying that for a given input graph, the embeddings generated by the surrogate should closely resemble those of the target.
GrOVe leverages this similarity in embeddings as its primary mechanism for ownership verification. The overall verification process is structured to be robust against various forms of adversarial behavior, including false accusations.
1. Model Registration:
The first crucial step is model registration. When a model owner (e.g., a company deploying a GNN service) finishes training their GNN, they must generate a cryptographic commitment (C) of their model. This commitment is essentially a unique fingerprint of the model's state, often a cryptographic hash of its parameters, architecture, or a combination thereof. The key property is that even minor changes to the model should result in a different commitment. Along with C, the owner must obtain a secure timestamp from a trusted third party. Both C and the timestamp are then registered with the third-party verifier. This ensures irrefutable proof of when the model was first deployed and its initial state.
2. Verification Process:
When an adversarial accuser claims that a suspect model (owned by an adversarial responder) was stolen from their target model, the third-party verifier initiates a multi-step process:
- Consistency Check: The verifier first checks if the target and suspect models are consistent with their registered commitments. This ensures that the models being evaluated are indeed the ones registered.
- Timestamp Verification: The verifier then consults the secure timestamps. It must be confirmed that the target model was registered before the suspect model. This temporal check is vital to prevent false accusations, as it establishes the chronological order of model creation.
- Verification Data Sampling: The verifier independently samples a verification data set from the same distribution as the target model's original training data. This data set is distinct from the target's training data but statistically similar, ensuring that the verification is fair and unbiased. Crucially, the verifier does not need access to the exact original training data, maintaining the black-box assumption.
- Model Query and Embedding Generation: The verifier then queries both the target GNN and the suspect GNN using the sampled verification data set. Both models output a set of node embeddings for each input graph.
- Verification Algorithm Execution: The generated embeddings from both models are then fed into GrOVe's core verification algorithm, which is a similarity classifier.
3. The Similarity Classifier (Verification Algorithm):
The central component of GrOVe is a specifically trained similarity classifier. Its objective is to determine whether a pair of embeddings (one from the target, one from the suspect) are "close" (indicating theft) or "far" (indicating independent training).
- Training the Similarity Classifier:
- Data Generation: To train this classifier, three types of GNNs are involved: a Target GNN, a Surrogate GNN (extracted from the target), and an Independent GNN (trained from scratch with the same architecture but different random initialization).
- Positive Samples: For each node in a given graph, embeddings are generated by both the Target GNN and the Surrogate GNN. A distance vector is computed between corresponding pairs of these embeddings (e.g., using L1 or L2 distance, or cosine similarity). These distance vectors, representing "close" embeddings, form the positive samples for the similarity classifier.
- Negative Samples: Similarly, embeddings are generated by the Target GNN and the Independent GNN. Distance vectors between these corresponding pairs form the negative samples, representing "far" embeddings.
- The similarity classifier is then trained on this dataset of positive and negative distance vectors to distinguish between "close" and "far" embedding relationships.
- Using the Similarity Classifier for Verification:
- During actual verification, the verifier obtains embeddings from the suspect GNN and the target GNN for the sampled verification graph.
- For each node, a distance vector is calculated between its embedding from the suspect model and its embedding from the target model.
- These distance vectors are then passed to the pre-trained similarity classifier.
- The classifier outputs a binary prediction for each pair: '1' if the embeddings are deemed "too close" (indicating a surrogate relationship) or '0' if they are "too far" (indicating an independent relationship).
- A threshold is applied to the proportion of '1' outputs. If a certain proportion (e.g., 0.5, as used in the paper) of the embedding pairs are classified as "too close," the verifier concludes that the suspect model is indeed a surrogate model extracted from the target. Otherwise, it is deemed an independent model. The authors found that a threshold of 0.5 worked effectively without requiring extensive optimization.
This detailed, multi-stage process, from registration to the core similarity classification, provides a robust framework for establishing and verifying ownership of GNNs.
Demo / Proof of Concept
▶ Watch: Model registration with cryptographic commitments (5:50)
While the talk did not feature a live, interactive demonstration of GrOVe in action, the researchers presented a comprehensive experimental setup and results that served as a compelling proof of concept for their methodology. The efficacy of GrOVe was rigorously tested under various conditions, validating its core principles.
The initial intuition and a crucial preliminary proof of concept were provided by t-SNE visualization. This technique was used to project the high-dimensional GNN embeddings into a 2D space, allowing for visual inspection. The results clearly showed that embeddings generated by a surrogate model for a given input fully overlapped with those from the target model, forming a single, coherent cluster. In stark contrast, embeddings from an independently trained GNN (with identical architecture but different random initialization) formed a distinct, separate cluster. This visual evidence provided strong initial validation that embedding similarity could indeed act as a reliable fingerprint for GNN ownership.
The subsequent experimental evaluation provided a more quantitative and robust proof of concept. The experimental setup involved:
- Metrics: Evaluating surrogate model accuracy (to ensure the extracted models were indeed functional replicas), false positive rate (FPR) (proportion of independent models misclassified as surrogates), and false negative rate (FNR) (proportion of surrogate models misclassified as independent).
- Similarity Classifier Training: The similarity classifier was trained using embeddings derived from Type 1 model extraction attacks as positive samples and embeddings from independently trained models as negative samples.
- Testing: To ensure unbiased evaluation, the similarity classifier was then tested on additional independent and surrogate models, generated with different random initializations than those used for training.
The results, as detailed in the "Key Findings" section, demonstrated GrOVe's effectiveness across multiple datasets, its robustness against evasion attempts like double extraction, and its adaptability to pruning through data augmentation. For instance, the consistent achievement of 0% FNR for both Type 1 and Type 2 attacks, coupled with very low FPRs (2.2% for Amazon, 3.4% for ACM, and 0% for others), serves as a powerful empirical demonstration of GrOVe's capabilities. Furthermore, the successful countermeasure against pruning, by augmenting the similarity classifier's training data with outputs from pruned models, illustrated the system's practical resilience. The efficiency benchmarks, showing classifier training in under 3 hours on an A100 GPU for the largest dataset, further underscore the practicality of this proof of concept for real-world deployment.
Defensive Implications
▶ Watch: Proposed verification algorithm using similarity classifier (7:20)
The GrOVe framework offers several critical defensive implications for organizations and individuals developing and deploying Graph Neural Networks:
- Proactive Ownership Registration: Model owners should proactively adopt a model registration process for all proprietary GNNs. This involves generating a cryptographic commitment (e.g., a hash of model weights or architecture) and obtaining a secure timestamp upon deployment. This foundational step provides irrefutable proof of ownership and the model's initial state, serving as the first line of defense against intellectual property theft claims.
- Integrate GrOVe Verification into IP Protection Strategy: Organizations should integrate GrOVe's embedding-based ownership verification as a standard procedure for intellectual property protection of their GNNs. This allows a trusted third-party verifier to ascertain the provenance of suspect models, especially when dealing with potential infringements or licensing disputes.
- Monitor for Embedding Similarity: The core insight of GrOVe — that extracted models produce highly similar embeddings — suggests that monitoring the embedding space could be a valuable defensive strategy. While GrOVe is designed for black-box verification, internal monitoring for unexpected embedding similarities among different deployed models could potentially flag suspicious activity or internal misuse.
- Awareness of Evasion Techniques and Countermeasures: Defenders must be aware of potential adversarial evasion techniques. GrOVe specifically addressed double extraction (which it effectively catches without modification) and pruning. For pruning, the research highlights the importance of data augmentation for the similarity classifier. This implies that as new evasion techniques emerge, the GrOVe framework can be adapted by incorporating examples of such evasions into the similarity classifier's training data, enhancing its robustness over time.
- Leverage Black-Box Verification: The black-box nature of GrOVe's verification process is a significant advantage. It allows a verifier to operate without needing access to the target model's sensitive training data or internal architecture details, only requiring prediction API access and a representative verification dataset. This makes it suitable for scenarios where intellectual property must be protected without revealing trade secrets.
- Adaptability to Other ML Models: While GrOVe is tailored for GNNs, the underlying principle of using embedding similarity as a fingerprint might be adaptable to other types of machine learning models that produce rich, high-dimensional representations (embeddings) of their inputs. This opens avenues for broader application of similar ownership verification techniques across the ML landscape.
By implementing these defensive strategies, GNN developers and deployers can significantly enhance the security posture of their valuable models against sophisticated model extraction threats.
Key Takeaways
- Graph Neural Networks (GNNs) are highly vulnerable to model extraction attacks, where adversaries can create functional replicas (surrogate models) by interacting with a target GNN's prediction API.
- GrOVe proposes an effective and efficient framework for ownership verification of GNNs by leveraging the inherent similarity of node embeddings generated by target and extracted surrogate models.
- The system includes a robust model registration process using cryptographic commitments and secure timestamps to establish verifiable ownership and prevent false accusations.
- GrOVe's core is a similarity classifier trained to distinguish between "close" embedding pairs (from target-surrogate) and "far" embedding pairs (from target-independent), achieving a 0% false negative rate and very low false positive rates in experiments.
- GrOVe is robust against advanced evasion techniques like double extraction and can be made robust against pruning through strategic data augmentation of the similarity classifier's training data.
- The framework is computationally efficient, with the similarity classifier training in less than 3 hours on an A100 GPU for large datasets, making it practical for real-world deployment.
About the Speaker(s)
The talk on GrOVe was presented by Asim Waheed, with contributions from co-authors Vasisht Duddu and N. Asokan. All three researchers are affiliated with the Secure Systems Group, indicating their expertise and focus on security aspects within complex computational systems, particularly in the realm of machine learning and graph-based applications. Their work demonstrates a commitment to addressing critical security challenges in state-of-the-art technologies like Graph Neural Networks.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research delivers a robust, practical solution to a critical problem: GNN intellectual property theft via model extraction. Leveraging embedding similarity as a fingerprint, GrOVe achieves a critical 0% false negative rate against known extraction attacks, making it a highly effective defensive mechanism. This isn't just theory; it's a well-engineered countermeasure with verifiable results.
Heather Calloway (CISO) — STRONG ACCEPT
This research on GrOVe offers a robust, actionable framework for verifying ownership of Graph Neural Networks, a critical step in protecting valuable AI intellectual property. Its focus on embedding similarity and a formal registration process provides a clear mechanism for institutional accountability against model extraction, directly addressing a significant business risk.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024