It's Simplex! Disaggregating Measures to Improve Certified Robustness

Andrew C. Cullen, Paul Montague, Shijie Liu, Sarah M. Erfani, Benjamin I.P. Rubinstein

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5

Overview

In the realm of machine learning security, adversarial examples pose a significant threat, capable of subtly altering input data to mislead classification models. While numerous reactive defenses have emerged, they often fall short against novel attack vectors, providing no generalizable security guarantees. This critical gap motivates the pursuit of certified robustness, a proactive approach that mathematically guarantees a lower bound on the distance to any potential adversarial example, thereby establishing a measurable radius of security around a given input.

Watch on YouTube

Visual summary for It's Simplex! Disaggregating Measures to Improve Certified Robustness by Andrew C. Cullen, Paul Montague, Shijie Liu, Sarah M. Erfani, Benjamin I.P. Rubinstein
Visual summary for It's Simplex! Disaggregating Measures to Improve Certified Robustness by Andrew C. Cullen, Paul Montague, Shijie Liu, Sarah M. Erfani, Benjamin I.P. Rubinstein

Key moments

  1. 2:00 Motivation for Certified Robustness against Adversarial Examples
  2. 2:40 Key advantages of certified robustness certificates
  3. 4:00 Understanding Randomized Smoothing for certified robustness
  4. 5:30 Reasons for Randomized Smoothing's popularity
  5. 5:40 Overview of Lacoa, Lee, and Cohen methods
  6. 7:50 Research question: Improving Lacoa's certificates with new DP advances

It's Simplex! Disaggregating Measures to Improve Certified Robustness

Speakers: Andrew C. Cullen, University of Melbourne; Paul Montague, DST Group; Shijie Liu, University of Melbourne; Sarah M. Erfani, University of Melbourne; Benjamin I.P. Rubinstein, University of Melbourne

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=LoxXBrBriQs

Overview

In the realm of machine learning security, adversarial examples pose a significant threat, capable of subtly altering input data to mislead classification models. While numerous reactive defenses have emerged, they often fall short against novel attack vectors, providing no generalizable security guarantees. This critical gap motivates the pursuit of certified robustness, a proactive approach that mathematically guarantees a lower bound on the distance to any potential adversarial example, thereby establishing a measurable radius of security around a given input.

This talk, "It's Simplex! Disaggregating Measures to Improve Certified Robustness," presented by Dr. Andrew Cullen and his collaborators from the University of Melbourne and DST Group, delves into advancements in this crucial area. The research introduces an improved certification scheme based on randomized smoothing and a novel analytical tool, the Simplex of Possible Outputs, to better understand and enhance certified robustness. The work not only pushes the boundaries of provable security for machine learning models but also offers a disaggregated view into the performance drivers of various certification techniques.

The core contribution lies in demonstrating how a refined understanding of differential privacy, combined with an intelligent ensembling of existing and new certification methods, can yield significantly stronger and more computationally efficient robustness guarantees. Furthermore, the Simplex of Possible Outputs provides a powerful diagnostic lens, allowing researchers and practitioners to pinpoint why certain samples exhibit varying levels of robustness and how model, data, or training influences these outcomes. This research is pivotal for developing more resilient AI systems that can withstand the ever-evolving landscape of adversarial attacks.

Background

▶ Watch: Motivation for Certified Robustness against Adversarial Examples (2:00)

The genesis of this research lies in the persistent challenge of adversarial examples, which are carefully perturbed inputs designed to trick machine learning classifiers. An attacker aims to find an altered input X' that, when fed to a classifier f, produces a different class prediction than the original input X, while minimizing the distance between X and X'. This minimal distance is crucial because it serves as a proxy for the detectability of the attack, either by automated systems or human observation. Most adversarial examples are generated using gradient descent techniques, making them fast, easy to execute, and resource-efficient.

In response, the security community has developed numerous defenses. However, these defenses are largely reactive and attack-specific. They aim to prevent particular attack types or increase their computational, financial, or time cost. The fundamental limitation is their lack of generalizability; a motivated adversary merely needs to discover an undefended threat vector to bypass these safeguards entirely. Such defenses offer no inherent guarantees of security.

This inherent weakness of reactive defenses has spurred interest in certified robustness. Instead of merely responding to known attacks, certified robustness aims to prove that no possible adversarial example can exist within a calculable finite radius around a given input. This approach offers two significant advantages: first, it actively forbids the "worst" adversarial examples—those that are infinitesimally close to the original input and thus extremely difficult to detect. Second, these robustness certificates can be used to rank samples by their inherent risk. Samples with larger certificates are provably more robust, requiring less scrutiny, while those with smaller certificates indicate higher potential risk, possibly necessitating manual intervention or further analysis.

A prominent mechanism for constructing these certificates is randomized smoothing. This technique involves taking an input sample, adding Gaussian noise with a fixed standard deviation, and then passing this perturbed sample through the model not once, but thousands or even tens of thousands of times. An expectation of the model's performance is then constructed based on the average predictions from these noisy inputs. While this process is inherently slow and slightly non-deterministic due to the Gaussian noise, its advantages are substantial: it is easily parallelizable, applicable to any blackbox model f, and imposes no additional limitations on the hardware beyond what is needed to run the model itself. This contrasts sharply with other certification methods that often demand specific GPU resources or impose significant architectural constraints on the target model, making randomized smoothing a highly popular and practical choice.

Within randomized smoothing, several key certification techniques have been proposed. Lacoa's approach, one of the earliest, leveraged differential privacy and was applicable to both softmax (average of class probabilities) and argmax (most frequently predicted class) style expectations. Lee later introduced an improvement using Renyi Divergence, though their method was specific to argmax expectations. Further advancements came from Cohen, who utilized the Neyman-Pearson Lemma to achieve even tighter guarantees, again exclusively for argmax-style expectations. The speakers observed that Lacoa's differential privacy mechanism, being general, could potentially be improved, especially given recent work by Bal et al. demonstrating how to tighten differential privacy guarantees specifically for systems involving Gaussian noise. This observation formed the core motivation for their investigation into developing a new, more robust certification scheme.

Key Findings

▶ Watch: Understanding Randomized Smoothing for certified robustness (4:00)

The research presented in "It's Simplex!" delivers several pivotal findings that significantly advance the field of certified robustness:

  1. Uniformly Improved Certification Scheme: The authors successfully developed a new certification scheme that leverages recent advancements in understanding differential privacy for Gaussian noise-based systems. This novel approach uniformly improves upon Lacoa's existing method for both softmax and argmax style expectations. Furthermore, it demonstrates superior performance over Lee's and Cohen's approaches for smaller adversarial radii, as measured by certified accuracy. This means models can achieve stronger, provable security guarantees across a broader range of attack scenarios.
  1. Efficient Ensemble Certification: A crucial insight of this work is the realization that when multiple independent robustness certificates are available for a given sample, the largest of these radii represents the guaranteed true robustness. By constructing an ensemble that simply takes the maximum of several certification techniques (including their newly proposed method, Lee's, and Cohen's), the researchers demonstrated a best-in-class performance. Critically, they found that the computational cost of creating such an ensemble is "infinitesimally small" compared to the cost of generating a single certificate, because the expensive expectations (derived from thousands of noisy model inferences) can be reused across all individual certification calculations. This provides a practical and efficient path to significantly higher certified robustness.
  1. The Simplex of Possible Outputs as a Diagnostic Tool: The paper introduces a novel visualization and analytical framework called the Simplex of Possible Outputs. This tool plots the relative performance of different certification techniques across the permissible output space of a classifier, specifically focusing on the two highest class expectations (e0 and e1). This allows for a disaggregated understanding of why certain techniques perform better for specific model outputs. By mapping the actual outputs of models (e.g., ResNet on ImageNet) onto this simplex, the tool reveals the influence of data, model architecture, and training processes on certification performance. It provides an unprecedented lens for diagnosing and understanding the drivers behind varying robustness levels, moving beyond aggregate metrics to provide actionable insights for model improvement.

Technical Deep Dive

▶ Watch: Reasons for Randomized Smoothing's popularity (5:30)

The technical foundation of "It's Simplex!" builds upon the established framework of certified robustness through randomized smoothing, while introducing a significant refinement to the underlying mathematical guarantees and a novel analytical tool.

At its core, randomized smoothing operates by transforming a base classifier f into a smoothed classifier g. For any input X, g(X) predicts the class c that f predicts most often when X is perturbed by Gaussian noise ε ~ N(0, σ^2I). The strength of this approach lies in the fact that the probability of g(X) changing its prediction under an L2 adversarial perturbation can be bounded. Specifically, if g(X) predicts class c_A and there exists an adversarial input X' that causes g(X') to predict c_B (where c_A ≠ c_B), then X' must be at least a certain distance away from X. This minimum distance is the certified radius R.

The construction of these certificates involves estimating the probabilities P(f(X+ε) = c) for each class c. This is done by passing X perturbed by thousands or tens of thousands of randomly sampled Gaussian noise vectors through the base classifier f and counting the predictions. These counts are then used to form expectations over the class outputs.

Existing randomized smoothing techniques, such as Lacoa's, Lee's, and Cohen's, differ in how they derive the certified radius from these expectations.

  • Lacoa's method, the earliest, employs principles of differential privacy to construct its certificates. It is notable for its versatility, applying to both softmax-style expectations (where the average of softmax outputs is considered) and argmax-style expectations (where the most frequently predicted class from noisy samples determines the expectation).
  • Lee's approach refined the certification for argmax-style expectations by utilizing the Renyi Divergence, leading to tighter bounds than Lacoa for certain scenarios.
  • Cohen's method further improved upon Lee's, again for argmax-style expectations, by applying the Neyman-Pearson Lemma. This technique is known for yielding some of the strongest certificates, particularly for larger radii.

The key technical innovation of "It's Simplex!" stems from a critical observation about Lacoa's differential privacy-based mechanism. Recent research by Bal et al. demonstrated methods to tighten the guarantees provided by differential privacy when the underlying noise mechanism is Gaussian. By incorporating these advancements, Cullen et al. were able to construct a new certification scheme. This scheme involves solving a maximization problem derived from the tightened differential privacy bounds. While this adds an extra step compared to standard approaches, the authors emphasize that the optimization surface is "quite regular" and the computational cost of finding the global maximum is "insignificant relative to the computational cost of constructing the expectations." This new scheme uniformly improves upon Lacoa's technique for both softmax and argmax expectations and shows promising results against Lee's and Cohen's for smaller radii.

To evaluate performance, the standard metric of certified accuracy is used. This measures the proportion of samples that possess a certified radius greater than a given L2 perturbation radius. Higher certified accuracy curves indicate better robustness. The results show that the new technique (referred to as "our technique" in the talk) uniformly outperforms Lacoa for softmax and surpasses Lee's and Cohen's for argmax up to a radius of approximately 1.0. Beyond this point, Cohen's approach tends to dominate.

A significant practical contribution is the concept of ensemble certification. The fundamental principle is that if multiple valid certified radii are computed for a single sample, the largest of these radii is the most robust guarantee. Therefore, an ensemble simply takes max(R_new, R_Lee, R_Cohen). The crucial insight here is the reuse of expectations. The most computationally expensive part of randomized smoothing is the repeated inference of the model with noisy inputs to generate the underlying class expectations. Once these expectations are computed, deriving the certified radius using Lacoa, Lee, Cohen, or the new technique involves relatively minor additional calculations. The authors demonstrate that the computational cost of running an ensemble is "infinitesimally small" compared to the cost of a single certification, as the dominant cost of sampling remains unchanged. This allows defenders to achieve a "best-in-class" certified accuracy curve, leveraging the strengths of each individual technique (e.g., the new technique for smaller radii and Cohen's for larger radii) without a proportional increase in computational overhead.

The most novel technical contribution is the Simplex of Possible Outputs. This is a visualization tool designed to disaggregate and understand the relative performance of different certification techniques. In a classification model with N classes, the output expectations (probabilities) for each class sum to one, forming an N-1 dimensional simplex. Since all current certification techniques primarily rely on the two highest class expectations (e0 and e1) to compute the radius, the Simplex of Possible Outputs can be effectively visualized in a 2D space defined by e0 and e1. By systematically evaluating the performance of different certification techniques across this (e0, e1) space, the researchers can precisely map regions where one technique significantly outperforms others. For example, for softmax, their approach can yield certificates over five times larger than Lacoa's in certain regions. For argmax, the simplex clearly delineates regions where their technique, Lee's, or Cohen's is dominant. When actual model outputs (e.g., from a ResNet model on ImageNet data) are plotted onto this simplex, it reveals the distribution of samples and which certification technique is most effective for different types of outputs. This provides an invaluable diagnostic tool for understanding the interplay between data characteristics, model behavior, and the efficacy of various certification methods, enabling a deeper understanding of the drivers of certified robustness.

Demo / Proof of Concept

▶ Watch: Overview of Lacoa, Lee, and Cohen methods (5:40)

While the talk did not feature a live, interactive demonstration in the traditional sense, the speakers presented compelling empirical evidence and visualizations that served as a robust proof of concept for their findings. These demonstrations effectively illustrated the improved performance of their new certification scheme, the benefits of ensembling, and the analytical power of the Simplex of Possible Outputs.

The core of the empirical proof was presented through certified accuracy curves. These plots graphically compared the proportion of samples that could be certified robust against an L2 perturbation of a given radius. The blue line, representing Lacoa's original softmax certification, was clearly outperformed by the dashed lines representing the new softmax certification (labeled "our technique"). For argmax-style certifications, the orange line (their technique) demonstrated superior performance over Lee's and Cohen's approaches up to a radius of approximately 1.0. Beyond this point, Cohen's method showed dominance. Crucially, the "multinomial" line, representing the ensemble approach (taking the maximum of all available certifications), consistently achieved the highest certified accuracy across the entire range of radii, effectively capturing the best performance from each individual technique.

Further quantitative evidence included tables comparing the median certified radius and the proportion of samples for which each technique yielded the largest certified radius. These statistics clearly showed that as the level of additive noise increased, their technique's median certified radius surpassed others, and it produced the largest radius for a greater proportion of samples. The ensemble, as expected, consistently delivered the largest median certified radius and certified more samples than any individual technique.

A critical aspect of the proof of concept was the computational cost analysis. A plot illustrating the computational cost per sample against the number of samples demonstrated an "incredibly small difference" between running a single certification and running the ensemble. This difference was shown to collapse as the number of samples increased, emphatically confirming that the dominant cost lies in the initial sampling process (generating expectations) and not in the subsequent calculation of multiple certificates from those shared expectations.

Finally, the Simplex of Possible Outputs served as a powerful visual proof of concept. The speakers showed plots of this simplex, illustrating the exact relative performance between their approach and Lacoa's for softmax, where their technique produced certifications up to five times larger. For argmax, the simplex was divided into distinct regions, visually demonstrating where their technique, Lee's approach, or Cohen's approach was dominant based on the e0 and e1 class expectations. By overlaying actual output points from an ImageNet dataset passed through a ResNet model using randomized smoothing onto this simplex, the researchers provided concrete evidence of how different samples fall into regions where specific certification techniques are most effective. This visualization not only validated their analytical tool but also provided a tangible way to understand the disaggregated influences of data and model on certification performance.

Defensive Implications

▶ Watch: Research question: Improving Lacoa's certificates with new DP advances (7:50)

The advancements presented in "It's Simplex!" carry significant implications for defenders seeking to build more robust and trustworthy machine learning systems. These findings offer practical strategies and diagnostic tools to enhance the security posture against adversarial examples.

Firstly, the development of a uniformly improved certification scheme means that defenders can now achieve stronger certified robustness guarantees for their models. By integrating the new differentially private certification method, especially for softmax-style outputs or for argmax outputs at smaller perturbation radii, practitioners can obtain higher certified radii. This translates directly into a larger provable safe zone around inputs, making models demonstrably more resilient to a wider range of adversarial attacks.

Secondly, the concept of efficient ensemble certification offers a highly practical and impactful defensive strategy. Defenders no longer need to choose a single "best" certification technique, which might perform optimally only in certain scenarios. Instead, by reusing the computationally expensive expectations generated during randomized smoothing, they can calculate multiple certificates (e.g., from the new method, Lee's, and Cohen's) with "infinitesimally small" additional cost and then simply take the maximum certified radius. This strategy ensures best-in-class performance across the entire spectrum of possible adversarial radii, maximizing the provable security for their models without incurring prohibitive computational overhead. This efficiency is critical for deploying certified models in real-world applications where computational resources are often constrained.

Thirdly, the Simplex of Possible Outputs provides an invaluable diagnostic and analytical tool for defenders. This visualization allows security researchers and ML engineers to move beyond aggregate metrics like certified accuracy and delve into the granular reasons behind a model's robustness (or lack thereof) for specific inputs. By plotting model outputs onto the simplex, defenders can identify regions where samples are consistently less robust, or where a particular certification technique is most effective. This disaggregated view helps in understanding the drivers of certification performance, including data influences, model architecture effects, and training process impacts. Such insights can guide targeted improvements, for example, by identifying problematic data subpopulations that might require augmentation or by suggesting model architectural changes that shift outputs into regions of higher certifiability.

Finally, this work reinforces the paradigm shift from reactive, attack-specific defenses to proactive, generalizable robustness guarantees. By providing mathematically proven lower bounds on the distance to adversarial examples, certified robustness offers a fundamentally stronger form of security. The improvements in efficiency and analytical depth offered by "It's Simplex!" make certified robustness more accessible and actionable for defenders, enabling them to build AI systems that are not just patched against known threats, but are inherently more robust and trustworthy from the ground up. This capability is crucial for critical applications where the integrity and reliability of AI predictions are paramount.

Key Takeaways

  • A new certification scheme, based on tightening differential privacy for Gaussian noise, uniformly improves upon Lacoa's approach for both softmax and argmax expectations, and outperforms Lee's and Cohen's for smaller adversarial radii.
  • Ensembling multiple randomized smoothing certification techniques (e.g., the new method, Lee's, and Cohen's) yields "best-in-class" certified robustness by simply taking the maximum radius, providing stronger guarantees.
  • The computational cost of ensembling is negligible because the expensive expectation calculations from randomized smoothing can be reused across all individual certification techniques, making it an efficient strategy.
  • The Simplex of Possible Outputs is a novel visualization and analytical tool that disaggregates and explains the relative performance of different certification techniques, helping to identify how data, model, and training influence robustness.
  • This research enhances the practicality and effectiveness of certified robustness, moving beyond reactive defenses to provide more efficient, provable, and diagnostically insightful security guarantees against adversarial examples.

About the Speaker(s)

The research presented in "It's Simplex! Disaggregating Measures to Improve Certified Robustness" was a collaborative effort by a team of academics and defense researchers. The talk was primarily delivered by Dr. Andrew Cullen from the University of Melbourne, Australia. His co-authors include Paul Montague from the DST Group in Adelaide, Australia, and Shijie Liu, Sarah M. Erfani, and Benjamin I.P. Rubinstein, all from the School of Computing and Information Systems at the University of Melbourne.

Their work was supported by various grants, including a DOD NGTF, an ARC DECR Grant, and computing resources supplied in part thanks to an ARC Leaf Grant. The team's broader research interests include other aspects of certified robustness, as highlighted by their related works such as "Double Bubble Toil and Trouble: Enhancing Certified Robustness Through Transitivity" (focused on exploiting transitive properties for improved certificates) and "AT2 Certifications: Robustness Certificates Yield Better Adversarial Examples" (which surprisingly explores how adversarial attacks can leverage robustness certificates to become more effective). This demonstrates a comprehensive engagement with the complexities of adversarial machine learning and the pursuit of robust AI systems.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research delivers a significant leap in certified robustness for ML, introducing a new, more efficient certification scheme and an ensemble approach that leverages existing techniques for best-in-class performance. The novel Simplex of Possible Outputs provides an invaluable diagnostic tool, moving beyond aggregate metrics to offer granular insights into model robustness. This is essential work for anyone building provably secure AI systems.

Heather Calloway (CISO) — STRONG ACCEPT

This research significantly advances certified robustness for machine learning, offering provable security guarantees against adversarial attacks. The introduction of an efficient ensemble certification method and the diagnostic Simplex of Possible Outputs provides security leaders with practical, measurable ways to manage AI risk and improve model resilience. It moves the conversation beyond reactive defenses to proactive, quantifiable security.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024