Automated Synthesis of Effect Graph Policies for Microservice-Aware Stateful System Call Specialization

William Blair, Frederico Araujo, Teryl Taylor, Jiyong Jang

IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 5

Overview

In an era dominated by cloud-native applications and microservice architectures, securing the underlying infrastructure against sophisticated attacks remains a paramount challenge for cloud operators. This talk, presented by William Blair and his collaborators Frederico Araujo, Teryl Taylor, and Jiyong Jang, from the IBM Thomas J. Watson Research Center, introduces a novel approach to enhance microservice security through Automated Synthesis of Effect Graph Policies for Microservice-Aware Stateful System Call Specialization. The core problem addressed is the inherent vulnerability of microservices, which often run arbitrary customer code within operating system containers, exposing cloud environments to significant risk if a container is compromised.

Watch on YouTube

Visual summary for Automated Synthesis of Effect Graph Policies for Microservice-Aware Stateful System Call Specialization by William Blair, Frederico Araujo, Teryl Taylor, Jiyong Jang
Visual summary for Automated Synthesis of Effect Graph Policies for Microservice-Aware Stateful System Call Specialization by William Blair, Frederico Araujo, Teryl Taylor, Jiyong Jang

Key moments

  1. 0:00 Introduction and microservice security problem
  2. 1:00 Unique security challenges for cloud operators
  3. 2:00 How system call specialization frameworks work
  4. 4:00 Expressiveness levels of existing security policies
  5. 5:00 Understanding the mimicry attack vulnerability
  6. 6:00 Introducing Micro-execution for stateful policies
  7. 7:00 Limitations of pure static analysis for policy inference

Automated Synthesis of Effect Graph Policies for Microservice-Aware Stateful System Call Specialization

Speakers: William Blair; Frederico Araujo; Teryl Taylor; Jiyong Jang

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=F4pOSYRphM0

Overview

In an era dominated by cloud-native applications and microservice architectures, securing the underlying infrastructure against sophisticated attacks remains a paramount challenge for cloud operators. This talk, presented by William Blair and his collaborators Frederico Araujo, Teryl Taylor, and Jiyong Jang, from the IBM Thomas J. Watson Research Center, introduces a novel approach to enhance microservice security through Automated Synthesis of Effect Graph Policies for Microservice-Aware Stateful System Call Specialization. The core problem addressed is the inherent vulnerability of microservices, which often run arbitrary customer code within operating system containers, exposing cloud environments to significant risk if a container is compromised.

Traditional system call specialization frameworks, while beneficial, frequently rely on stateless or overly permissive security policies. These policies, often derived from static analysis or superficial dynamic tracing, fail to capture the nuanced, stateful behavior of microservices configured for specific tasks. This oversight leaves a critical gap, enabling mimicry attacks where adversaries exploit legitimate, yet unused, system calls within a microservice's allowed set to achieve malicious goals. The presented research directly tackles this by proposing Effect Graphs as a more expressive, stateful security policy representation, capable of precisely defining and enforcing the expected system call sequences and concrete arguments for a given microservice configuration.

The significance of this work lies in its potential to dramatically improve the containment of compromised microservices. By moving beyond simple allow-lists to context-aware, stateful policies, MicrPolicyCraft, the framework introduced in the talk, offers a robust defense against advanced threats that evade conventional system call filters. This innovation is crucial for cloud operators who provide arbitrary computation, as it empowers them to automatically generate highly precise security policies without relying on customers to author them, thereby fortifying the security posture of large-scale cloud applications against both known and emergent attack vectors.

Background

▶ Watch: Introduction and microservice security problem (0:00)

The evolution of application deployment has seen a significant shift from monolithic applications to highly distributed microservice architectures, often running within operating system containers. While containers offer benefits like isolation and portability, they introduce unique security challenges. Cloud operators must allow customers to run arbitrary code, which means they cannot always predict the exact programs or configurations that will be executed. This unpredictability makes it difficult to author precise security policies manually, and relying on customers to do so is often impractical or unreliable.

One established approach to mitigate this risk is system call specialization. This technique aims to restrict a microservice's behavior by limiting the set of system calls it can issue to the kernel. At its core, system call specialization involves analyzing a container image – comprising the microservice entry point, configuration files, and application dependencies – to produce a security policy. This policy is then enforced at runtime by a reference monitor, such as a seccomp filter implemented using eBPF within the Linux kernel. If a microservice attempts to issue a system call not permitted by its assigned policy, the reference monitor terminates the service, preventing further malicious activity.

However, existing system call specialization frameworks exhibit varying levels of expressiveness, each with inherent limitations:

  1. Stateless Policies: The simplest form, treating policies as a "bag of system calls." Any system call not in this set triggers a violation. These policies are highly prone to over-approximation because they permit all system calls found in the binary, regardless of whether they are needed for a specific configuration.
  2. Resource-Limited Policies: Frameworks like AppArmor extend stateless policies by restricting allowed system calls to specific resources (e.g., files, network ports). While an improvement, they still lack behavioral context.
  3. Temporal Policies: More advanced policies partition an application's lifecycle into distinct phases (e.g., "boot" and "serve"), allowing different sets of system calls in each phase. This introduces a basic form of state, but the transitions are often coarse-grained and still vulnerable.

Despite these advancements, a significant vulnerability persists: mimicry attacks. In a mimicry attack, an adversary compromises a microservice but performs malicious actions using only system calls that are allowed by the existing security policy. For instance, a general-purpose web server configured only to serve static files might still have the execve system call present in its binary. A stateless policy would likely permit execve. If an attacker compromises the server, they could then use execve to spawn a shell or run arbitrary commands, all while operating within the "allowed" system call set, thus evading detection. This problem is particularly acute for general-purpose programs, which, by design, contain a broad range of functionalities and associated system calls, many of which are not required for a specific deployment configuration. The challenge, therefore, is to create security policies that are not only restricted to a specific set of system calls but also capture the sequence and concrete arguments of these calls, precisely reflecting the intended operational behavior of a microservice.

Key Findings

▶ Watch: How system call specialization frameworks work (2:00)

The research presents several key findings and contributions that significantly advance the state of microservice security:

  1. Introduction of Effect Graphs as Stateful Security Policies: The most fundamental contribution is the definition and formalization of Effect Graphs. These are novel graph-based data structures that represent a microservice's expected behavior by capturing not just individual system calls, but their sequences and the concrete arguments passed to them. This stateful representation overcomes the limitations of stateless or simple temporal policies, providing a much more precise and context-aware security model.
  2. MicrPolicyCraft Architecture for Automated Synthesis and Monitoring: The talk introduces MicrPolicyCraft, a comprehensive framework designed to both automatically synthesize Effect Graph policies and efficiently monitor microservice executions for violations. This end-to-end solution integrates binary analysis, micro execution, and a runtime policy monitor, providing a practical system for deploying these advanced security policies.
  3. Micro Execution for Precise Policy Generation: The core of policy synthesis relies on a technique called micro execution. This innovative approach combines the strengths of static analysis (for control flow recovery) and dynamic analysis (for observing concrete execution paths and arguments). By abstracting external library dependencies and leveraging container image context, micro execution can generate highly accurate and restricted effect graphs, avoiding the over-approximations common in purely static methods and the incompleteness of purely dynamic tracing.
  4. Efficient Runtime Monitoring with Security Automata: The Microservice-aware Policy Monitor (MPM), a component of MicrPolicyCraft, efficiently detects policy violations by treating effect graphs as security automata. This allows the MPM to rapidly check whether observed system call traces conform to the expected state transitions and argument constraints defined in the effect graph. Performance evaluations demonstrated that this monitoring can be done with minimal overhead, maintaining manageable resource utilization and not negatively impacting application request runtime, even under high load (e.g., an Nginx web server handling a million requests).
  5. Effective Coverage and Policy Conciseness: The evaluation of MicrPolicyCraft against real-world programs showed "good effect coverage," meaning the generated effect graphs accurately captured the intended behaviors. Crucially, the policy size (number of states in the effect graph) did not grow significantly even as the binaries increased in size, demonstrating that precise, configuration-specific policies can be maintained in a concise manner, avoiding the bloat of over-approximated policies.
  6. Protection for Entire Microservice Architectures: The framework extends its capabilities to secure complex, multi-service applications by composing individual microservice effect graphs into a distributed effect graph. This allows for the detection of policy violations that span across service boundaries, such as an attacker in one compromised service attempting to communicate with an unauthorized database in another.

Technical Deep Dive

▶ Watch: Expressiveness levels of existing security policies (4:00)

The technical foundation of this work rests on Effect Graphs, which are a novel, stateful representation of a microservice's expected system call behavior. An Effect Graph is a directed graph where:

  • Nodes represent specific system call terms within the program's Control Flow Graph (CFG).
  • Edges between nodes represent the sequence of system calls.
  • Each node is annotated with concrete arguments observed during micro execution. This allows for fine-grained control, such as restricting a read system call to a specific file path or a connect call to a particular IP address and port.
  • A syntactic sugar is introduced where a comma denotes sequences of system call nodes occurring on the same resource, simplifying representation.

The generation of these precise effect graphs is achieved through micro execution, a hybrid analysis technique designed to overcome the limitations of purely static or dynamic approaches:

  1. Static Analysis Limitations: While static analysis can recover the CFG and identify potential system call paths, it often leads to over-approximation. Resolving destinations for indirect jumps for all configurations is challenging, and inferring concrete inputs for every execution path is often impossible without runtime context.
  2. Dynamic Analysis Limitations: Pure dynamic tracing captures concrete system call sequences but struggles to map these back to specific parts of the original program's CFG. It's difficult to determine, for example, if a series of read calls are part of a loop or distinct operations without the program's structural context.
  3. Micro Execution's Hybrid Approach: Micro execution combines these by performing a lightweight dynamic analysis within a controlled environment. It starts with a static analysis to recover the CFG of the microservice's entry point. Then, it symbolically executes the program, but instead of fully symbolic execution, it interacts with the actual layered file system of the container image and leverages test inputs (e.g., network inputs) to observe concrete system call arguments and sequences. This process is facilitated by an effect tracking plugin implemented on top of an existing micro execution framework called Primis, which was improved to abstract external library dependencies through an Application Binary Interface (ABI). This allows Primis to execute parts of the binary, observe actual system calls, and record them as an Effect Graph, restricted to the specific configuration and inputs provided.

The synthesized effect graphs are then utilized by the MicrPolicyCraft architecture for runtime monitoring:

  • Policy Generation Framework: Takes the container image (layered file systems) and test inputs. The binary entry point is passed to a binary analysis platform to lift its CFG. This CFG, along with configuration files and library dependencies, is fed into the micro execution framework (Primis with the effect tracking plugin). The output is a stateful Effect Graph policy for each microservice.
  • Runtime Monitoring: Microservices run normally within a cloud orchestration engine. If a microservice is compromised, the CIS flow Telemetry system generates a lightweight system call trace (CIS flow Trace).
  • Microservice-aware Policy Monitor (MPM): Implemented as a plugin within the CIS flow Telemetry framework. The MPM treats each Effect Graph as a security automaton. It attempts to "advance" this automaton using the actions (system calls and arguments) observed in the CIS flow Trace. If an observed action causes the automaton to become "stuck" (i.e., there is no valid transition for that system call/argument sequence), the MPM detects a policy violation and raises an alert to the operator.
  • Symbolic Constraints: To enhance flexibility, effect graphs can incorporate symbolic constraints. For instance, a policy might allow connections to .example.com or access to files within /var/log/ using regular expressions, rather than requiring explicit enumeration of every possible hostname or file path.

This robust framework addresses several research challenges: accurately modeling external libraries, recovering precise CFGs and system call arguments, scaling the analysis to real-world programs, and efficiently monitoring for policy violations at the computing edge without incurring significant performance overhead.

Demo / Proof of Concept

▶ Watch: Introducing Micro-execution for stateful policies (6:00)

The talk effectively demonstrates the capabilities of MicrPolicyCraft through illustrative examples and performance evaluations, showcasing both policy generation and runtime detection.

Policy Synthesis Example:

The process of generating an effect graph is exemplified by modeling a simple server that periodically reads and writes data to disk. During micro execution with Primis, the framework observes a sequence of read and write system calls. Because these occur in a loop, the resulting effect graph concisely represents this alternating sequence. Importantly, each node in the graph is annotated with the concrete arguments observed for these calls (e.g., the specific database file being accessed). This yields a security policy G that precisely encodes the server's behavior: "reads and writes some data to a particular database file in alternating sequence." The evaluation highlights that this approach achieves "good effect coverage" for real-world artifacts (measured by the ratio of effects in the graph to total effectful terms in the binary) and maintains "concise policy size" (number of states) even as binary sizes increase, ensuring practicality.

Runtime Policy Violation Detection:

A key demonstration involves an Nginx web server. An effect graph is generated for Nginx, representing its legitimate operations. The scenario postulates an adversary compromising the Nginx instance and attempting to establish a connection to an external command-and-control (C2) network. The CIS flow Telemetry system captures this malicious connect system call in a trace. The Microservice-aware Policy Monitor (MPM), treating the Nginx effect graph as a security automaton, attempts to process the CIS flow trace. When it encounters the connect call to the C2 server, it finds no valid transition in the automaton for this specific action, causing the automaton to become "stuck." This immediately triggers a policy violation alert, demonstrating the MPM's ability to detect mimicry attacks that would bypass simpler, stateless policies.

Performance Evaluation:

The talk presents compelling performance metrics to underscore the efficiency of the MPM. In a scenario where an Nginx web server acts as a configuration proxy, handling a million requests for a backend application server, the MPM's overhead is negligible. The Nginx's resource utilization remains manageable, and, critically, the request runtime is "not negatively impacted" by the continuous security monitoring. The MPM itself is lightweight, consuming "up to two cores and less than 50 megabytes of RAM" on each physical computing node. Furthermore, a single MPM instance can concurrently monitor "multiple effect graphs" against the same CIS flow Telemetry trace without inhibiting overall microservice architecture operation.

Protection of Entire Microservice Architectures:

The framework's capability to secure complex, multi-service applications is illustrated with a hotel reservation microservice architecture. Imagine a "search" microservice within this architecture becomes compromised. An adversary, instead of following the intended communication links, attempts to directly access a MongoDB database to disclose user location history. The distributed effect graph, formed by the composition of all individual microservice effect graphs, defines the legitimate inter-service communication patterns. The MPM observes the unauthorized direct mongodb access, determines it does not conform to the distributed effect graph's model, and raises a violation. This demonstrates the framework's ability to detect policy violations across an entire architecture, enforcing intended communication boundaries.

Defensive Implications

▶ Watch: Limitations of pure static analysis for policy inference (7:00)

The insights presented by MicrPolicyCraft offer profound implications for security practitioners and cloud operators aiming to bolster their defenses against sophisticated attacks in microservice environments.

  1. Shift from Stateless to Stateful Policies: Defenders must recognize the inherent limitations of traditional stateless or resource-limited system call policies. Relying solely on a "bag of allowed system calls" leaves organizations vulnerable to mimicry attacks, where legitimate system calls are abused for malicious purposes. The imperative is to adopt stateful security policies that consider the sequence and concrete arguments of system calls, reflecting the true intended behavior of an application.
  2. Automated Policy Generation is Key: Manually crafting precise security policies for every microservice, especially those built from general-purpose binaries, is impractical and error-prone. Organizations should seek or develop tools that can automatically synthesize highly specific policies based on observed configurations and expected behaviors. Frameworks like MicrPolicyCraft, which combine static and dynamic analysis via micro execution, represent the future of policy generation, reducing the burden on security teams while increasing policy accuracy.
  3. Enhanced Runtime Monitoring: Beyond simply checking if a system call is allowed, defenders need to implement microservice-aware policy monitors that treat policies as security automata. This enables the detection of anomalous sequences of system calls or calls with unforeseen arguments, even if the individual calls themselves are technically allowed. Integrating such monitors with existing CIS flow Telemetry systems can provide real-time, high-fidelity alerts for subtle attack patterns.
  4. Containment of General-Purpose Binaries: The problem of over-permissive policies for general-purpose programs (e.g., an Nginx server, a database client) is directly addressed. By generating configuration-specific effect graphs, defenders can significantly reduce the attack surface of these ubiquitous components, preventing exploitation of functionalities not required for their specific deployment. This means a web server configured for static files will only be allowed to perform file I/O operations relevant to serving those files, not execve or arbitrary network connections.
  5. Securing Inter-Service Communication: Modern applications are composed of many communicating microservices. Defenders must extend their security posture beyond individual service boundaries to encompass the entire microservice architecture. Tools that can compose individual policies into a distributed effect graph are essential for detecting unauthorized communication flows between services, preventing lateral movement or data exfiltration attempts that bypass intended application logic.
  6. Leverage Open Source Tools: The work highlights that the tool and its dependencies are open source. This encourages security researchers and practitioners to explore, adapt, and integrate these advanced techniques into their existing security stacks, fostering collaborative development and adoption of more robust microservice security paradigms.

Key Takeaways

  • Mimicry attacks exploit over-permissive, stateless system call policies in microservices, enabling adversaries to use allowed system calls for malicious purposes.
  • Effect Graphs are a novel, stateful security policy representation that captures expected system call sequences and concrete arguments, providing precise behavioral enforcement.
  • MicrPolicyCraft is an end-to-end framework for automatically synthesizing Effect Graph policies using micro execution (a hybrid static/dynamic analysis) and monitoring them at runtime.
  • The Microservice-aware Policy Monitor (MPM) treats effect graphs as security automata to efficiently detect policy violations in real-time with minimal performance overhead.
  • MicrPolicyCraft enables the generation of highly specific policies for general-purpose binaries, significantly reducing the attack surface by restricting behavior to a microservice's exact configuration.
  • The framework can secure entire microservice architectures by composing individual effect graphs, detecting unauthorized inter-service communication and containing sophisticated threats.

About the Speaker(s)

The talk was presented by William Blair, who conducted this work during an internship with his talented supervisors at the IBM Thomas J. Watson Research Center. His collaborators on this research include Frederico Araujo, Teryl Taylor, and Jiyong Jang. The research focuses on advanced security techniques for cloud-native applications, specifically addressing the challenges of securing microservices through automated policy synthesis and stateful system call specialization. Their affiliation with the IBM Thomas J. Watson Research Center suggests a background in cutting-edge computer science research, particularly in areas of systems security, software analysis, and cloud computing.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research introduces Effect Graphs and MicrPolicyCraft, a novel framework for automated, stateful system call specialization that directly counters mimicry attacks in microservices. By combining hybrid analysis with security automata, it generates precise, configuration-specific policies, significantly advancing cloud-native security. This is a critical development for containing compromised services in complex, distributed environments.

Heather Calloway (CISO) — STRONG ACCEPT

This research offers a critical advancement for microservice security, directly addressing the significant operational risk posed by over-permissive system call policies. Automated synthesis of stateful Effect Graphs provides a precise, scalable defense against mimicry attacks, translating directly into enhanced containment and reduced attack surface for cloud-native applications.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024