From Principle to Practice: Vertical Data Minimization for Machine Learning

Robin Staab, Nikola Jovanovic, Mislav Balunovic, Martin Vechev

IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 6

Overview

This talk, presented by Robin Staab and his colleagues from the SML Lab at ETH Zurich, introduces a novel approach to data privacy known as Vertical Data Minimization (VDM) for machine learning. The core premise of VDM is to reduce the granularity of collected personal data attributes by generalizing them, rather than simply collecting fewer data points (horizontal minimization). This methodology directly addresses the stringent requirements of privacy regulations like Europe's GDPR Article 5(c) and the US AI Bill of Rights, which mandate that personal data collection be "adequate, relevant, and limited to what is necessary." The work specifically targets the untrusted collector setting, focusing on data collected during model deployment and making minimal assumptions about client capabilities.

Watch on YouTube

Visual summary for From Principle to Practice: Vertical Data Minimization for Machine Learning by Robin Staab, Nikola Jovanovic, Mislav Balunovic, Martin Vechev
Visual summary for From Principle to Practice: Vertical Data Minimization for Machine Learning by Robin Staab, Nikola Jovanovic, Mislav Balunovic, Martin Vechev

Key moments

  1. 0:00 Introduction and definition of data minimization
  2. 2:00 Vertical data minimization: concept and example
  3. 2:50 Motivation: Fines and limitations of existing PETs
  4. 3:20 Introducing the adversarial setting for VDM
  5. 4:25 VDM's unique role compared to other PETs
  6. 5:00 The typical workflow of vertical data minimization
  7. 6:20 Algorithms and baselines for learning generalizations
  8. 7:40 Proposed: Privacy-aware Tree Short Pad Algorithm (P-Jinny)

From Principle to Practice: Vertical Data Minimization for Machine Learning

Speakers: Robin Staab, Nikola Jovanovic, Mislav Balunovic, Martin Vechev

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=Jq_zLptqoMY

Overview

This talk, presented by Robin Staab and his colleagues from the SML Lab at ETH Zurich, introduces a novel approach to data privacy known as Vertical Data Minimization (VDM) for machine learning. The core premise of VDM is to reduce the granularity of collected personal data attributes by generalizing them, rather than simply collecting fewer data points (horizontal minimization). This methodology directly addresses the stringent requirements of privacy regulations like Europe's GDPR Article 5(c) and the US AI Bill of Rights, which mandate that personal data collection be "adequate, relevant, and limited to what is necessary." The work specifically targets the untrusted collector setting, focusing on data collected during model deployment and making minimal assumptions about client capabilities.

The significance of this research lies in its practical applicability to real-world scenarios where organizations collect sensitive user data for machine learning inference. By minimizing data vertically, the proposed methods aim to significantly reduce the risk of privacy breaches and associated fines, which can reach millions of euros. Unlike many existing privacy-enhancing technologies (PETs) that focus on model training or modify how data is collected, VDM adapts what kind of data is collected, making it a complementary and easily combinable solution. The talk not only defines VDM but also introduces a comprehensive adversarial evaluation framework and presents state-of-the-art algorithms, with their Privacy Aware Tree Short Path (PAD) method consistently outperforming all others.

Background

▶ Watch: Introduction and definition of data minimization (0:00)

The concept of data minimization is not new, deeply rooted in privacy regulations globally. GDPR's Article 5(c) explicitly states that personal data should be "adequate, relevant, and limited to what is necessary in relation to the purposes for which it is processed." More recently, the US AI Bill of Rights echoed this sentiment, emphasizing that AI systems should collect only data "strictly necessary for the specific context." These regulations highlight a critical gap in current practices, particularly in scenarios involving an untrusted data collector where personal data is often collected during model deployment. Traditional data minimization often refers to horizontal data minimization, where fewer data points (e.g., fewer users) are collected. While effective for model training, this approach is impractical for model inference, where specific individuals still need to be processed.

This work addresses the limitations of horizontal minimization by proposing vertical data minimization. Instead of reducing the number of data points, VDM focuses on reducing the resolution or specificity of individual attributes. For example, instead of collecting exact ages, VDM might generalize them into broader age ranges (e.g., "18-25", "26-35"). This directly reduces the amount of personal data while aiming to maintain utility for downstream machine learning tasks.

Existing privacy-enhancing technologies (PETs) often concentrate on securing data during model training (e.g., federated learning, differential privacy) or modifying the how of data collection (e.g., homomorphic encryption). However, many PETs either impose strong client assumptions, do not provide guarantees during the initial collection phase, or are not feasible for inference in an untrusted collector setting. The authors argue that VDM is largely orthogonal to these existing PETs. It changes what data is collected, rather than how it's collected or how a model is trained, meaning it can be easily combined with other privacy measures to create a more robust privacy posture. Prior work in this space, such as privacy-preserving data publishing (PPDP), often focuses on protecting against re-identification when publishing datasets. In contrast, the VDM adversarial setting specifically targets the reconstruction of generalized attributes directly, a more challenging and relevant threat model for dynamic inference scenarios.

Key Findings

▶ Watch: Motivation: Fines and limitations of existing PETs (2:50)

The research yields several significant findings that underscore the efficacy and importance of vertical data minimization:

  1. Superior Performance of PAD: The Privacy Aware Tree Short Path (PAD) algorithm consistently outperforms all other baseline and novel VDM methods across various datasets and adversarial scenarios. It achieves a remarkable balance between maintaining high utility for downstream ML tasks and providing strong privacy guarantees against adversarial reconstruction. On the ACs employment dataset, PAD achieved almost the same utility as non-minimized data while sacrificing only minimal amounts of privacy.
  2. Significant Privacy-Utility Tradeoff Gains: The study demonstrates that there is substantial room for improvement in the privacy-utility tradeoff for data collection. By employing VDM, it's possible to achieve robust privacy protection without a drastic reduction in model accuracy. The Pareto curves clearly illustrate that advanced VDM methods, particularly PAD, push the boundary of what's achievable in this tradeoff space, moving closer to the ideal of high utility and high privacy.
  3. Adversaries Can Be More Accurate on Certain Individuals: Even with VDM applied, adversaries focusing on high-certainty reconstructions (A2 adversary) can make noticeably more accurate predictions for specific individuals where their confidence is high. This highlights the ongoing need for robust privacy measures and careful evaluation, even when data is generalized.
  4. Resilience Against Additional Adversarial Knowledge: While adversaries with additional knowledge (A3-A5) can improve their reconstruction capabilities, data minimization still provides a significant protective effect. Even in the worst-case scenario where an adversary knows all but one personal attribute, VDM demonstrably helps to obscure the remaining target attribute.
  5. Vulnerability to Multi-Breach Scenarios: The multi-breach scenario (A6) reveals that if an adversary gains access to data under multiple different generalizations (e.g., a data collector switches generalization strategies over time), the adversarial error for a joint breach can be noticeably lower than the minimum error from individual breaches. This emphasizes the importance of consistent and robust generalization strategies over time.
  6. Protection Against Linkability and Singling Out: VDM methods, including PAD, offer valuable protection against modern privacy threats like linking initially disjoint datasets (A7) and uniquely identifying individuals within a minimized dataset (A8), which are crucial considerations informed by data protection working party documents.
  7. Stability and Robustness: The PAD algorithm demonstrates strong individual attribute privacy protection and exhibits stability even under reasonable distribution shifts across different data splits, as shown on the ACs employment dataset.

Technical Deep Dive

▶ Watch: VDM's unique role compared to other PETs (4:25)

The core objective of vertical data minimization is to learn a generalization function G that transforms original, full-resolution data S_orig into a minimized version G(S_orig). This process must adhere to three critical requirements:

  1. Global: The generalization G must be consistent across all clients.
  2. Single-dimensional: Each attribute is generalized independently.
  3. Strict: G must form a strict partition of the input attribute domain, ensuring every value maps to exactly one generalized value.

The workflow for VDM involves training G on a small set of full-resolution records. For adversarial scenarios, it's assumed an adversary can observe these data pairs to learn reconstruction algorithms. After G is selected and trained, only minimized data is ever collected for both model training and deployment, ensuring compliance with data minimization regulations.

The talk explores several algorithms for learning these generalizations:

Baselines

  • Uniform Splitting: Divides input ranges into equally sized bins.
  • Feature Selection Mechanisms: Common ML techniques that select relevant features, which can be seen as a form of VDM if less precise features are chosen.
  • Goldin et al. (A): A prior work based on tree pruning algorithms.
  • Iterative: A more involved baseline using dynamic programming to find optimal generalization boundaries.

Novel Minimizers

The research introduces more sophisticated algorithms, starting with neural network-based approaches:

  • Neural Minimizers: These methods leverage learnable neural network representations to achieve both high utility and high privacy. Inspired by Generative Adversarial Networks (GANs), they employ a joint objective function with one term optimizing for utility (e.g., downstream task accuracy) and another for privacy (e.g., adversarial reconstruction error).
  • Mutual Information Minimizer: This variant replaces the adversarial privacy objective of the neural minimizer with one based on mutual information. The goal here is to minimize the mutual information between the original and generalized attributes, thereby reducing the information leakage.

Privacy Aware Tree Short Path (PAD)

The most significant contribution is the Privacy Aware Tree Short Path (PAD) algorithm, which consistently outperforms other methods. PAD builds upon classical tree learning algorithms but integrates privacy considerations directly into the tree construction process.

  1. Decision Tree Construction: Similar to standard decision trees, PAD builds a tree by recursively splitting internal nodes based on a single attribute.
  2. Privacy Aware Gini (P-Gini) Criterion: The core innovation lies in adapting the classical Gini splitting criterion. The standard Gini impurity measures how often a randomly chosen element from the set would be incorrectly labeled if it were randomly labeled according to the distribution of labels in the subset. P-Gini extends this by adding an additional privacy term. This privacy term, inspired by fair representation learning, intuitively aims to make the distribution of all personal attributes as uniform as possible across any split in the decision tree. This makes it harder for an adversary to infer original values based on the split.
  3. Alpha Parameter: A crucial parameter, Alpha, allows for a continuous tradeoff between privacy and utility during the tree construction. A higher Alpha prioritizes privacy, potentially at the cost of utility, and vice-versa.
  4. Generalization Extraction: Once the tree is built, the generalization is derived by projecting all decision boundaries from the inner split nodes onto the respective attribute. For example, if a tree splits an "age" attribute at 2 and 7, the generalized age categories would be 0-1, 2-6, and 7+.

Adversarial Evaluation Scenarios (A1-A8)

To rigorously evaluate the privacy threat, the authors propose eight diverse and regulation-aligned evaluation scenarios, directly informed by the EU's data protection working party:

  • A1 (General Reconstruction): The adversary attempts to reconstruct personal attributes from minimized data, having learned a reconstruction algorithm by observing full-resolution data pairs used to train G.
  • A2 (High Certainty Reconstruction): Similar to A1, but the adversary's performance is evaluated only on predictions where they are highly certain, reflecting a focus on individual privacy violations.
  • A3-A5 (Additional Knowledge Adversaries): These scenarios model adversaries with varying levels of auxiliary knowledge:
  • A3: Adversary has access to all non-personal attributes.
  • A4: Worst-case scenario where the adversary has access to all attributes except the current target personal attribute.
  • A5: An intermediate case between A3 and A4.
  • A6 (Multi-Breach Scenario): Models the case where an adversary has access to data generated under multiple different generalizations (e.g., if a data collector changes their VDM strategy over time).
  • A7 (Linkability): Evaluates how well minimized data can be used to link two previously non-linked datasets, a common re-identification threat.
  • A8 (Singling Out): Checks whether individual entries in the minimized dataset can be uniquely identified, another critical re-identification concern.

These scenarios provide a comprehensive framework for assessing the robustness of VDM methods against a wide range of realistic adversarial capabilities and objectives.

Demo / Proof of Concept

▶ Watch: The typical workflow of vertical data minimization (5:00)

The talk presents compelling empirical results to demonstrate the effectiveness of VDM methods, particularly PAD, using the ACs employment dataset. The evaluation focuses on the Pareto curve, plotting classification error (utility) on the X-axis against adversarial error (privacy) on the Y-axis.

Key points on the Pareto curve illustrate the extremes:

  • Fully Generalized (top-right): Represents a scenario where no useful data is collected (all data is the same). Both the adversary and the utility inference model predict only statistical majority classes, leading to high adversarial error (good privacy) but also high classification error (low utility).
  • No Generalization (bottom-left): Represents the case where full, non-minimized data is always sent. This provides an upper bound on utility (low classification error) but offers minimal privacy (low adversarial error).

The optimal goal is to be in the top-left of this graph, achieving high utility with high privacy.

The results consistently show:

  • PAD's Dominance: When plotting the Pareto curves for all evaluated methods and their hyperparameters, PAD consistently dominates all other approaches across the utility-privacy tradeoff. It achieves utility levels almost identical to non-minimized data while introducing only small sacrifices in privacy, positioning it very close to the ideal top-left corner. This observation holds true across all evaluated datasets.
  • A2 Adversary (High Certainty): The evaluation with the A2 adversary reveals that, in practice, adversaries can make noticeably more accurate predictions on individuals for whom they have higher certainty. This underscores the nuanced nature of privacy protection and the importance of considering targeted attacks.
  • A3-A5 Adversaries (Additional Knowledge): While additional knowledge inherently helps adversaries, the minimization provided by VDM methods still significantly hinders their reconstruction efforts. Even in the challenging A4 scenario, where all but one personal attribute is known, VDM offers a noticeable privacy boost.
  • A6 Adversary (Multi-Breach): The multi-breach scenario demonstrates that combining data from different generalizations can lead to a lower adversarial error than the minimum error from any single breach. This highlights a potential vulnerability that needs to be considered in deployment strategies.
  • A7 (Linkability) and A8 (Singling Out): The talk confirms that generalized data, when processed through VDM, offers insights into its resilience against these critical threats. The methods provide a degree of protection against linking disjoint datasets and uniquely identifying individuals, aligning with the concerns raised by data protection working parties.

Further evaluations presented in the paper, briefly mentioned in the talk, indicate that PAD leads to strong individual attribute privacy protection despite its joint objective and exhibits stability to reasonable distribution shifts, as demonstrated on five data set splits of the ACs employment dataset. These results collectively validate VDM as a powerful and practical approach to enhancing data privacy in machine learning contexts.

Defensive Implications

▶ Watch: Proposed: Privacy-aware Tree Short Pad Algorithm (P-Jinny) (7:40)

The findings from this research offer critical insights for organizations and data custodians striving to comply with data minimization principles and protect user privacy. The primary defensive implication is the strategic adoption of vertical data minimization as a fundamental component of their data collection and processing pipelines, especially in scenarios involving untrusted data collectors or model inference.

Defenders should consider the following actionable steps:

  1. Integrate VDM Early: Implement VDM at the point of data collection, before full-resolution personal data ever leaves the client or is stored in a potentially untrusted environment. This proactive approach aligns directly with regulations requiring data to be "limited to what is necessary" from the outset.
  2. Prioritize PAD or Similar Algorithms: Organizations should evaluate and adopt state-of-the-art VDM algorithms like Privacy Aware Tree Short Path (PAD). The consistent superior performance of PAD in balancing utility and privacy makes it a strong candidate for practical deployment.
  3. Conduct Adversarial Privacy Evaluations: Beyond simple utility checks, organizations must rigorously evaluate their chosen VDM strategies against diverse adversarial models, similar to the A1-A8 scenarios proposed. This includes assessing resilience against:
  • Reconstruction attacks (A1, A2, A3-A5) with varying levels of adversary knowledge.
  • Multi-breach scenarios (A6), which necessitate consistent and robust generalization strategies over time to prevent adversaries from combining information from different data versions.
  • Linkability (A7) and singling out (A8) threats, which address modern re-identification risks.
  1. Understand the Privacy-Utility Tradeoff: Organizations must consciously define their acceptable privacy-utility tradeoff based on the sensitivity of the data and regulatory requirements. Tools like the Alpha parameter in P-Gini allow for fine-tuning this balance, enabling organizations to meet specific compliance thresholds.
  2. Complement Existing PETs: Recognize that VDM is largely orthogonal to other privacy-enhancing technologies (PETs). It can and should be combined with other measures like differential privacy (for training), secure multi-party computation, or homomorphic encryption, to create a layered and more comprehensive privacy defense strategy. VDM addresses what data is collected, complementing PETs that focus on how data is processed or stored.
  3. Monitor and Adapt: The threat landscape evolves. Organizations should continuously monitor the effectiveness of their VDM implementations against new attack vectors and distribution shifts, as PAD has shown stability in this regard. Regular re-evaluation and adaptation of generalization strategies will be crucial for long-term privacy protection.
  4. Document Generalization Strategies: For audit and compliance purposes, organizations should meticulously document their chosen generalization functions, the rationale behind them, and the results of their privacy evaluations. This demonstrates due diligence in data minimization efforts.

By integrating these defensive implications, organizations can move from principle to practice, effectively implementing vertical data minimization to significantly enhance data privacy, reduce regulatory risk, and build greater trust with their users.

Key Takeaways

  • Vertical Data Minimization (VDM) is crucial for inference: Unlike horizontal minimization, VDM reduces the granularity of individual attributes (e.g., age ranges instead of exact age) and is essential for maintaining utility during machine learning inference while complying with privacy regulations like GDPR and the AI Bill of Rights.
  • VDM is orthogonal and complementary to other PETs: VDM modifies what data is collected, making it easily combinable with existing privacy-enhancing technologies that focus on how data is processed or secured during training, offering a layered defense.
  • PAD algorithm sets a new state-of-the-art: The Privacy Aware Tree Short Path (PAD) algorithm consistently outperforms other methods by effectively balancing utility for downstream ML tasks and privacy against adversarial reconstruction, demonstrating significant gains in the privacy-utility tradeoff.
  • Comprehensive adversarial evaluation is vital: The research introduces eight regulation-aligned adversarial scenarios (A1-A8) to rigorously assess VDM methods against various threats, including reconstruction, multi-breach attacks, linkability, and singling out.
  • Adversaries can exploit certainties and multiple breaches: Even with VDM, adversaries can make more accurate predictions on individuals where they are highly certain, and access to data from multiple generalization strategies (multi-breach) can significantly lower adversarial error.
  • VDM provides robust protection even with auxiliary knowledge: While additional adversarial knowledge can improve reconstruction, VDM demonstrably helps protect personal attributes, even in challenging scenarios where most other attributes are known.

About the Speaker(s)

The research presented in this talk is a collaborative effort by Robin Staab, Nikola Jovanovic, Mislav Balunovic, and Martin Vechev. All authors are affiliated with the SML Lab at ETH Zurich, Switzerland. Robin Staab, who presented the paper, is a key contributor to this work, representing the team's expertise in machine learning and security. Martin Vechev is their advisor, guiding the research within the SML Lab. Their collective work focuses on advancing the theoretical and practical aspects of secure and private machine learning systems, with a particular emphasis on addressing real-world challenges posed by data privacy regulations.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research on Vertical Data Minimization (VDM) for ML inference is a critical, high-impact contribution. It introduces a genuinely novel approach to data minimization, distinct from existing PETs, and provides a state-of-the-art algorithm (PAD) along with a robust adversarial evaluation framework. This directly addresses pressing regulatory and privacy challenges, offering actionable solutions for real-world deployments.

Heather Calloway (CISO) — STRONG ACCEPT

This research provides a crucial, actionable framework for Vertical Data Minimization, directly addressing regulatory compliance and reducing real-world business risk in machine learning deployments. The PAD algorithm offers an evidence-backed method to balance privacy with utility, giving security leaders concrete steps to operationalize data minimization and strengthen institutional accountability.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024