General-Purpose f-DP Estimation and Auditing in a Black-Box Setting

Önder Askin (University Boom)

34th USENIX Security Symposium (USENIX Security '25) · Day 2 · Privacy 1: Differential Privacy and Audit

Overview

This talk, presented by Önder Askin from the University of Bochum, introduces novel methods for estimating and auditing f-differential privacy (f-DP) in a black-box setting. f-DP is a powerful and increasingly popular variant of differential privacy, offering easily interpretable privacy guarantees and suitability for complex privacy-preserving system design, particularly in auditing private machine learning models. The core challenge addressed is how to verify the privacy claims of a mechanism when its internal code or structure is unknown, a common scenario in real-world deployments involving proprietary software or third-party libraries.

Watch on YouTube · Slides

Visual summary for General-Purpose f-DP Estimation and Auditing in a Black-Box Setting by Önder Askin
Visual summary for General-Purpose f-DP Estimation and Auditing in a Black-Box Setting by Önder Askin

Key moments

  1. 0:00 Introduction to F-Differential Privacy estimation and auditing
  2. 2:20 Formalizing F-DP: Adversary's guessing game and hypothesis testing
  3. 4:16 Optimal test for error combination: The Likelihood Ratio Test
  4. 5:39 Understanding the trade-off curve T in F-Differential Privacy
  5. 6:52 Reasons for F-Differential Privacy's increasing popularity and utility
  6. 8:00 Estimating F-DP privacy parameters in a black-box scenario
  7. 8:47 Using output samples to approximate densities for F-DP estimation

General-Purpose f-DP Estimation and Auditing in a Black-Box Setting

Speakers: Önder Askin

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=I4tjMo7esI4

Overview

This talk, presented by Önder Askin from the University of Bochum, introduces novel methods for estimating and auditing f-differential privacy (f-DP) in a black-box setting. f-DP is a powerful and increasingly popular variant of differential privacy, offering easily interpretable privacy guarantees and suitability for complex privacy-preserving system design, particularly in auditing private machine learning models. The core challenge addressed is how to verify the privacy claims of a mechanism when its internal code or structure is unknown, a common scenario in real-world deployments involving proprietary software or third-party libraries.

The research, a collaborative effort with Soladet Martin Duner Timka Yunlu Juve and Vasileis Seekers, provides a robust framework to tackle this problem. It devises a statistical approach to estimate the f-DP trade-off function from observed outputs, even without access to the mechanism's underlying probability distributions. Furthermore, the talk details a method to audit privacy claims with statistical significance, enabling reliable detection of privacy violations and exposing mechanisms that fail to meet their stated privacy guarantees.

This work holds significant importance for the security and privacy community. As differential privacy becomes a cornerstone for protecting sensitive data in machine learning and other data-driven applications, the ability to independently verify privacy assurances is paramount. The black-box nature of the proposed methods makes them directly applicable to a wide array of practical scenarios, empowering auditors, regulators, and users to hold data processors accountable for their privacy claims.

Background

▶ Watch: Introduction to F-Differential Privacy estimation and auditing (0:00)

Differential privacy (DP) is a rigorous mathematical framework for quantifying and guaranteeing privacy in data analysis. The standard definition, often referred to as (ε, δ)-differential privacy, considers a database D and a mechanism M that produces an output M(D). Privacy is assessed by comparing the output distributions of M on D and a "neighboring" database D', where D' differs from D by the information of exactly one individual. The core idea is that if an adversary cannot distinguish M(D) from M(D'), then the presence or absence of any single individual's data in D has minimal impact on the output, thereby ensuring privacy. The parameters ε and δ quantify this indistinguishability: smaller values typically indicate stronger privacy.

f-differential privacy (f-DP) refines this notion of indistinguishability by framing it through the lens of an optimal adversary playing a guessing game. Given an output x, the adversary attempts to determine whether x originated from database D or D'. This decision task can be formulated as a hypothesis testing problem: H0 (null hypothesis) states x came from D, and H1 (alternative hypothesis) states x came from D'. Any hypothesis test deployed by the adversary is subject to two types of errors:

  • Type I error (α): The probability of falsely rejecting H0 (i.e., concluding x came from D' when it actually came from D).
  • Type II error (β): The probability of falsely failing to reject H0 (i.e., concluding x came from D when it actually came from D').

The adversary's goal is to minimize these errors. For a fixed Type I error level α, the optimal test that achieves the smallest possible Type II error β is the Likelihood Ratio Test (LRT). The LRT compares the ratio of probability densities P(x) (from M(D)) and Q(x) (from M(D')) against a threshold. By varying α from 0 to 1, one can plot the corresponding minimal β values, resulting in a trade-off curve (T). This curve is the central object of interest in f-DP, as it encapsulates the fundamental limits of distinguishability between the outputs of M on D and D'.

A mechanism M is formally defined as f-DP for a function F if any trade-off curve T implied by two neighboring databases D and D' lies above F. Intuitively, a higher trade-off curve T indicates greater indistinguishability and thus stronger privacy. f-DP offers several advantages over (ε, δ)-DP: it provides easily interpretable privacy guarantees, is well-suited for reasoning about the composition of algorithms (combining multiple DP mechanisms while maintaining overall privacy), and facilitates the study of privacy amplification (how privacy can improve under certain conditions, like subsampling). Crucially, f-DP has gained traction as a robust framework for auditing the privacy properties of machine learning models.

Key Findings

▶ Watch: Optimal test for error combination: The Likelihood Ratio Test (4:16)

The core contributions of this research revolve around enabling practical, black-box estimation and auditing of f-differential privacy. The key findings include:

  1. Black-Box Estimation of f-DP Trade-off Functions: The researchers devised a novel approach to estimate the f-DP trade-off function (T) for a mechanism operating in a black-box setting. This means the method does not require access to the mechanism's internal code or the underlying probability distributions P and Q of its outputs. Instead, it relies solely on observed output samples generated by running the mechanism on a database D and its neighboring counterpart D'.
  1. Introduction of a Perturbed Likelihood Ratio Test: To overcome the practical difficulties of directly mimicking the optimal likelihood ratio test (LRT) with estimated densities (specifically, the challenge of the estimated density ratio rarely equaling an exact threshold for coin-toss decisions), a perturbed Likelihood Ratio Test was introduced. This perturbed test, while slightly modified, has a trade-off function (T_h) that converges uniformly to the true trade-off function (T) as a perturbation parameter h approaches zero. This theoretical guarantee underpins the validity of the estimation approach.
  1. Uniform Convergence of the Estimator: The proposed estimator for the perturbed trade-off function (T_h_hat), derived using Kernel Density Estimators (KDEs), demonstrates uniform convergence to the true trade-off function T as the sample size N grows. This is a critical theoretical result, indicating that with sufficient data, the estimated privacy level will accurately reflect the true privacy level of the mechanism. The talk illustrates this with examples showing how T_h_hat for N=5,000 samples closely overlaps the true T compared to N=200 samples.
  1. Statistically Significant Auditing Method: Beyond mere estimation, the research provides a method for auditing f-DP privacy claims with statistical significance. By focusing on single points of the trade-off function, the method leverages optimal classification theory to derive estimates with associated confidence intervals. These confidence intervals are then used to construct a confidence region around the estimated trade-off curve. If a claimed privacy level (T0) does not overlap with this confidence region, the claim can be reliably rejected with high probability, providing a robust, quantifiable basis for detecting privacy violations.
  1. Reliable Detection of Privacy Violations: The auditing framework moves beyond observing a "visible discrepancy" between claimed and actual privacy levels. By providing theoretical guarantees and confidence regions, the methods enable auditors to reliably infer and detect privacy violations, exposing flawed mechanisms with a controlled level of statistical error. This is crucial for building trust and accountability in privacy-preserving systems.

Technical Deep Dive

▶ Watch: Understanding the trade-off curve T in F-Differential Privacy (5:39)

The technical core of this work lies in overcoming the challenge of estimating the f-DP trade-off function T in a black-box setting. In such a scenario, an auditor has no access to the internal workings of the mechanism M, nor to the true probability density functions P and Q of its outputs on neighboring databases D and D', respectively. The only available information consists of output samples.

The first step involves generating these samples. The auditor runs the mechanism M repeatedly on D to obtain n samples x1, ..., xn corresponding to M(D). Similarly, M is run on a carefully chosen neighboring database D' to obtain n samples y1, ..., yn corresponding to M(D'). These samples form the empirical basis for the estimation.

The trade-off function T is fundamentally defined using the Likelihood Ratio Test (LRT), which requires knowledge of the densities P and Q. Since these are unknown, the research proposes using Kernel Density Estimators (KDEs) to approximate them from the collected samples. A KDE for P(x) (denoted P_hat(x)) is essentially a weighted sum of "density bumps" (kernel functions K) centered around each observation xi:

P_hat(x) = (1 / (n B)) sum_{i=1 to n} K((x - xi) / B)

where B is the bandwidth parameter that controls the smoothness of the estimate, and K is the kernel function (e.g., Gaussian kernel). A similar estimator Q_hat(x) is constructed for Q(x) using samples y1, ..., yn.

Plugging P_hat and Q_hat directly into the original LRT definition presents a practical hurdle. The optimal LRT rejects the null hypothesis H0 if the density quotient Q(x)/P(x) exceeds a threshold eta, or if it equals eta and a coin toss (with a specific probability lambda) yields "heads." In practice, with estimated densities, the ratio Q_hat(x)/P_hat(x) will rarely be exactly equal to eta. This makes accurately mimicking the coin toss decision difficult and introduces instability in the estimation.

To circumvent this, the authors introduce a perturbed Likelihood Ratio Test. This perturbed test rejects H0 if Q_hat(x)/P_hat(x) lies above eta + h * u, where h is a small positive constant and u is a random variable uniformly drawn from [-1/2, 1/2]. This perturbation effectively "smooths out" the exact equality condition, making the test amenable to estimation from samples. Crucially, the trade-off function T_h associated with this perturbed test is shown to converge uniformly to the true trade-off function T as h approaches zero.

With the perturbed LRT, an estimator for T_h (denoted T_h_hat) can be computed. The points on T_h_hat are calculated based on the chosen threshold eta. To obtain an estimate for the entire curve, these points are computed for a range of eta values. The talk illustrates this with the LLAS mechanism example: for N=200 samples, the T_h_hat curve (red) provides an approximation, but for N=5,000 samples, the T_h_hat curve (yellow) almost perfectly overlaps the true T (blue), demonstrating the uniform convergence of the estimator with increasing sample size. This performance, even with "moderate" sample sizes like N=5,000, is a significant achievement for black-box estimation.

For auditing privacy claims, the approach moves from estimating the entire curve to focusing on specific points. In f-DP, a larger trade-off curve signifies more privacy. If a mechanism claims privacy level T0 (a specific trade-off function), but its actual trade-off function T lies below T0, then a privacy violation has occurred. The auditing method aims to statistically detect this discrepancy.

The key insight for auditing is to leverage optimal classification theory by focusing on single points (alpha, beta) on the trade-off functions. This allows the construction of confidence intervals for the estimated alpha and beta values at specific points, given a confidence parameter gamma. These confidence intervals quantify the uncertainty around the estimates.

To construct a confidence region for auditing, the method first identifies the point (alpha_max, beta_max) where the distance between the estimated T_h_hat and the claimed T0 is largest. Around this point, a confidence region is constructed using the previously derived confidence intervals. The logic is as follows: the confidence region is designed to contain the expected values of the estimates with high probability. These expected values correspond to the true (alpha, beta) points of the mechanism. If the claimed privacy level T0 does not overlap with this confidence region, it implies that the true trade-off curve T (which is expected to lie within the confidence region) is indeed below T0, thereby indicating a statistically significant privacy violation. The talk demonstrates this with a "flawed Gaussian mechanism" example, where a visible gap between T_h_hat and T0, coupled with a non-overlapping confidence region, allows for the reliable rejection of the false privacy claim.

Demo / Proof of Concept

▶ Watch: Estimating F-DP privacy parameters in a black-box scenario (8:00)

While the talk did not feature a live, interactive "demo" in the traditional sense of a software demonstration, it presented compelling illustrations and numerical examples that served as a proof of concept for the proposed estimation and auditing methods. These examples effectively demonstrated the theoretical underpinnings and practical applicability of the research.

One key illustration involved the LLAS mechanism (Laplace mechanism for sensitivity 1, with additive noise), a common differentially private mechanism. The speaker presented a plot showing:

  1. The true trade-off function T (blue curve) for the LLAS mechanism, which represents the actual privacy afforded by the mechanism.
  2. An estimate T_h_hat (red curve) obtained using the perturbed likelihood ratio test with a moderate sample size of N=200. This visually demonstrated the initial approximation capability of the estimator.
  3. A more refined estimate T_h_hat (yellow curve) obtained with a larger sample size of N=5,000. This plot clearly showed how the yellow curve almost perfectly overlapped the true blue curve, providing strong visual evidence of the uniform convergence of the estimator with increased sample size. This illustrated that even with relatively moderate sample sizes, the black-box estimation method could accurately infer the f-DP properties.

For the auditing aspect, the talk used an example of a flawed Gaussian mechanism. Here, the speaker presented:

  1. A privacy claim T0 (a specific trade-off function) that was deliberately set to be "too optimistic," suggesting a higher level of privacy than the flawed mechanism actually provided.
  2. An estimate T_h_hat for the true privacy parameter of the flawed mechanism.
  3. A confidence region constructed around the point of largest discrepancy between T_h_hat and T0.

The visual proof of concept showed a clear distance between the estimated T_h_hat and the claimed T0. Crucially, the constructed confidence region did not overlap with the claimed T0. This non-overlap allowed the speaker to conclude with high probability that the privacy claim T0 was false, thereby identifying a privacy violation in the flawed Gaussian mechanism. These illustrations served to validate the practical utility of the proposed methods for both estimating f-DP curves and reliably auditing privacy claims in real-world scenarios.

Defensive Implications

▶ Watch: Using output samples to approximate densities for F-DP estimation (8:47)

The research presented on general-purpose f-DP estimation and auditing in a black-box setting offers several critical implications for security defenders and organizations deploying differentially private systems:

  1. Independent Verification of Privacy Claims: This work provides a powerful tool for organizations to independently verify the privacy guarantees of mechanisms, especially those developed by third parties or integrated as proprietary black-box components. Instead of solely relying on vendor claims or theoretical papers, defenders can now empirically audit whether a mechanism truly adheres to its f-DP specification.
  1. Detection of Privacy Violations in Production: The ability to audit mechanisms in a black-box setting is invaluable for detecting privacy violations in deployed systems. If a mechanism's implementation deviates from its specification, or if subtle bugs introduce privacy leaks, these methods can expose such flaws. This moves beyond pre-deployment static analysis or white-box testing to runtime verification, which is crucial for dynamic and evolving systems.
  1. Enhanced Accountability and Trust: By enabling statistically significant auditing, the research fosters greater accountability for privacy claims. Organizations can demonstrate due diligence by regularly auditing their DP mechanisms, building trust with users and complying with regulatory requirements that demand verifiable privacy assurances.
  1. Robustness Against Implementation Errors: Even well-intentioned developers can introduce errors that inadvertently weaken privacy. The proposed black-box auditing framework acts as a safety net, catching such implementation flaws that might not be apparent from code review alone. This is particularly relevant for complex machine learning models where privacy-preserving mechanisms can be intricate.
  1. Guidance for Mechanism Selection and Parameter Tuning: While primarily an auditing tool, the estimation method can also inform mechanism selection and parameter tuning. By observing the trade-off curves for different configurations or mechanisms, defenders can gain a clearer understanding of the actual privacy-utility balance achieved, guiding decisions on which mechanisms to deploy and with what parameters.
  1. Continuous Privacy Monitoring: The techniques can be integrated into continuous integration/continuous deployment (CI/CD) pipelines or ongoing monitoring systems. Regular black-box audits could be performed to ensure that privacy guarantees remain consistent as data evolves, models are updated, or underlying infrastructure changes. This helps prevent privacy degradation over time.
  1. Applicability to Machine Learning Auditing: Given f-DP's popularity in auditing private machine learning, these methods are directly applicable to verifying the privacy of differentially private machine learning models. This is particularly important for models trained on sensitive data, where privacy guarantees are paramount and often mandated. Defenders can use this to check if a private ML model truly satisfies its stated f-DP properties against various attacks.
  1. Understanding "False Optimism": The concept of a privacy claim (T0) being "too optimistic" is a key takeaway. Defenders should be aware that theoretical claims might not always translate perfectly to practice. This auditing framework provides a concrete way to identify and rectify such discrepancies, ensuring that users are not given a false sense of security regarding their data privacy.

Key Takeaways

  • f-differential privacy (f-DP) offers a robust and interpretable framework for quantifying privacy, especially valuable for auditing privacy-preserving machine learning algorithms and reasoning about algorithm composition.
  • The research introduces a novel, black-box approach for estimating the f-DP trade-off function, requiring only output samples from the mechanism and no access to its internal code or underlying probability distributions.
  • Kernel Density Estimators (KDEs) are used to approximate the unknown output densities, and a perturbed Likelihood Ratio Test addresses practical challenges in mimicking the optimal LRT, enabling robust estimation.
  • The proposed estimator demonstrates uniform convergence to the true f-DP trade-off function as the number of samples increases, providing strong theoretical guarantees for the accuracy of the estimation.
  • A statistically rigorous auditing method is presented, which constructs confidence regions around estimated privacy levels. This allows for the reliable and statistically significant rejection of false f-DP privacy claims with a controlled probability of error.
  • These methods empower auditors and defenders to verify privacy claims independently in real-world systems, enhancing accountability, detecting privacy violations, and building trust in differentially private deployments.

About the Speaker(s)

The talk was presented by Önder Askin from the University of Bochum. This research represents a collaborative effort, with significant contributions from his colleagues Soladet Martin Duner Timka Yunlu Juve and Vasileis Seekers. While specific titles or institutional affiliations for the collaborators were not detailed in the transcript, their involvement underscores the interdisciplinary nature and collective expertise behind this advanced work in differential privacy and security auditing.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid, technically grounded academic work that fills a real gap: black-box auditing of f-DP with statistical guarantees, directly applicable to the growing mess of 'we use differential privacy' claims from ML vendors and cloud providers. The perturbed LRT construction and uniform convergence result are non-trivial contributions, and the auditing framework with confidence regions is exactly what the space needs.

Heather Calloway (CISO) — WEAK

Technically serious work on black-box f-DP auditing with real theoretical contributions — but it never closes the gap between the math and the operator. The defensive implications section reads like a marketing appendix, not a deployment guide, and the accountability framing stays aspirational throughout.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)