Heimdall: Towards Risk-Aware Network Management Outsourcing

Yuejie Wang

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Network Security 1

Overview

In an era where operational efficiency and cost reduction drive business decisions, the outsourcing of IT services, particularly network management, has become a pervasive trend. This talk, "Heimdall: Towards Risk-Aware Network Management Outsourcing," presented by Yuejie Wang at the NDSS Symposium, addresses the critical security challenges inherent in this growing practice. The presentation introduces Heimdall, a novel framework designed to define, quantify, monitor, and respond to the risks associated with outsourcing network configuration troubleshooting. Given that the managed services market is projected to reach over $300 billion by 2025, with major players like Verizon, Fujitsu, and IBM offering such services, the security implications of granting third-party access to production networks are profound and increasingly complex.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction and problem: Risk-aware network management outsourcing
  2. 1:18 Defining HDA's goals and threat model
  3. 2:30 Proposed asset-based quantitative risk model
  4. 4:00 Overcoming limitations with data plane-leveraged dependency graphs
  5. 5:00 HDA's end-to-end workflow and system components
  6. 6:20 Observation 1: Operator access order affects consequence likelihood
  7. 7:40 Observation 2: Root cause estimation using historical data

Heimdall: Towards Risk-Aware Network Management Outsourcing

Speakers: Yuejie Wang

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=vZiFAWkve0

Overview

In an era where operational efficiency and cost reduction drive business decisions, the outsourcing of IT services, particularly network management, has become a pervasive trend. This talk, "Heimdall: Towards Risk-Aware Network Management Outsourcing," presented by Yuejie Wang at the NDSS Symposium, addresses the critical security challenges inherent in this growing practice. The presentation introduces Heimdall, a novel framework designed to define, quantify, monitor, and respond to the risks associated with outsourcing network configuration troubleshooting. Given that the managed services market is projected to reach over $300 billion by 2025, with major players like Verizon, Fujitsu, and IBM offering such services, the security implications of granting third-party access to production networks are profound and increasingly complex.

The core problem Heimdall aims to solve stems from the common workflow where enterprise administrators create tickets for network issues, which are then addressed by third-party technicians. These technicians often require direct, privileged access to the production network. The issuance of a single misconfigured or malicious command by such a technician can precipitate catastrophic outcomes, ranging from ransomware injection to large-scale service outages. Heimdall's objective is to transform this risky workflow by providing a risk-aware mechanism that quantifies the potential impact of network changes, guards the network through continuous risk monitoring, and enforces appropriate responses, thereby mitigating the significant security concerns associated with outsourced network management.

Background

▶ Watch: Introduction and problem: Risk-aware network management outsourcing (0:00)

The increasing adoption of network management outsourcing, driven primarily by cost efficiency, has inadvertently created a substantial attack surface. While enterprises benefit from specialized expertise and reduced operational overhead, they simultaneously introduce a new class of risk: the potential for compromise or error by an external entity with privileged access. The focus of Heimdall is specifically on network configuration troubleshooting, a frequent and inherently risky activity that often involves modifying critical network infrastructure.

Existing approaches to risk assessment in network environments typically associate risk with individual commands. For example, a shutdown command might be deemed riskier than a show command. However, this model suffers from several significant limitations. Firstly, it lacks flexibility; the same command can carry vastly different risks depending on the specific network configuration context (e.g., modifying a core OSPF area versus an aged one). Secondly, these models are often coarse-grained. As highlighted in the talk, Cisco IOS products, for instance, support only up to 16 privilege levels for commands, which is insufficient to capture the nuanced risk profiles of diverse network operations. Furthermore, current methods often rely on simple token matching on configuration files to identify dependencies. This can lead to an explosion of false positive dependency links, falsely linking unrelated configuration blocks (e.g., all interfaces within a router sharing an OSPF protocol token), resulting in a dense and unmanageable risk consequence graph. These shortcomings underscore the need for a more sophisticated, context-aware, and quantitative approach to risk assessment in outsourced network management.

Heimdall's underlying threat model acknowledges that while enterprise administrators are generally trustworthy, third-party technicians, despite their expertise in solving tickets, could also be the source of network incidents, either through inadvertent error or malicious intent. This necessitates a framework that can not only identify potential risks but also dynamically enforce policies to prevent their realization, without impeding the legitimate troubleshooting process.

Key Findings

▶ Watch: Proposed asset-based quantitative risk model (2:30)

Heimdall introduces a paradigm shift in managing outsourced network configuration troubleshooting by focusing on asset-based quantitative risk assessment rather than command-centric, coarse-grained models. The framework's key findings and contributions can be summarized as follows:

  1. Quantitative, Asset-Based Risk Definition: Heimdall defines the risk of a ticket (Risk(T)) as the sum of the conditional probabilities (P(S | T)) that an asset S can be affected during ticket resolution, multiplied by the value of that asset (S.value). This approach fundamentally links network operations to their potential impact on valuable enterprise assets, providing a more precise and business-relevant risk metric.
  1. Data Plane-Leveraged Risk Dependency Graph (RDG): Unlike traditional token-matching methods, Heimdall constructs its Risk Dependency Graph (RDG) by leveraging data plane information. This allows for the establishment of causal links between configuration blocks based on how routes are established, broadcasted, or learned, significantly reducing false positive dependencies and generating a more accurate consequence graph. The talk notes that pure token matching can generate significantly more false positive dependency links.
  1. Incorporation of Behavioral Factors: Heimdall enhances the accuracy of P(S | T) by integrating two critical behavioral observations:
  • Preference Order: The model accounts for the varying expertise of technicians by abstracting their diagnostic and resolution process as a preference order on configuration blocks within the RDG subgraph. An expert technician, for instance, can identify and prioritize the root cause block more quickly.
  • Root Cause Estimation: The framework incorporates root cause estimation, computed using historical statistics, to provide an educated guess about the likely source of a problem before it is fully resolved. This allows for proactive risk management based on predictive analytics.
  1. Efficient and Scalable Risk Monitoring and Response: Heimdall's centralized reference monitor can process commands quickly, incurring a negligible overhead of only four to five milliseconds per command. This efficiency ensures that the risk-aware privilege management workflow does not significantly impede troubleshooting. The system demonstrates that 86% of tasks incur less than 10% overhead on this workflow.
  1. Effective Risk Reduction: Through comprehensive expert validation involving 99 rounds of experiments over 41 hours, Heimdall proved capable of effectively reducing outsourcing risks. A notable finding was that 92% of tasks were solved without incurring any extra risks beyond those associated with the root cause block itself, highlighting its precision in containing potential damage.
  1. Scalability for Large Networks: The RDG construction and subsequent risk computations are designed to scale easily to hundreds of routers in large backbone networks. Critically, all computation time for these processes is proactive, occurring before a technician even begins solving a ticket, thus avoiding delays during critical troubleshooting phases.

Technical Deep Dive

▶ Watch: Overcoming limitations with data plane-leveraged dependency graphs (4:00)

Heimdall's architecture is composed of three primary components: Risk Definition, Risk Assessment, and Risk Monitoring and Response. Each component plays a crucial role in establishing a comprehensive, risk-aware framework for outsourced network management.

Risk Definition

The foundational principle of Heimdall's risk model is that the risk of an event is the probability of its occurrence multiplied by the value loss incurred by its consequence. To operationalize this, Heimdall introduces the concept of assets as instances with values that are of primary concern to an enterprise. These assets can include physical equipment (e.g., routers, switches) and software components.

A ticket in the system represents a network issue that needs resolution. The resolution of such a ticket can be projected onto a series of events that modify network configurations. For a given ticket T, Heimdall defines the conditional probability P(S | T) as the likelihood that an asset S will be affected during the resolution of T.

The aggregate risk of a ticket T is then formally defined by the sum of P(S | T) multiplied by the value of each affected asset S.value across all assets attached to the network:

Risk(T) = Σ (P(S | T) * S.value)

This asset-based approach moves beyond generic command risks to provide a quantitative measure directly tied to an enterprise's critical resources.

Risk Assessment Model

Accurately assessing P(S | T) is central to Heimdall. Existing token-matching approaches on configuration files are inherently flawed because they often falsely link unrelated configuration blocks, leading to an exponentially dense and difficult-to-reason-about risk consequence graph. Heimdall addresses this by constructing a Risk Dependency Graph (RDG) that leverages data plane information in addition to textual analysis.

The construction of the RDG is guided by a set of five rules designed to establish causal links between configuration blocks. The core idea is to identify when a route is established, broadcasted, or learned, rather than simply looking for shared tokens. For instance, if a change to Block A causes a route to be advertised, and Block B learns this route, a causal dependency is established between A and B. This method provides a far more precise understanding of how modifications propagate through the network.

Beyond the structural dependencies captured by the RDG, Heimdall incorporates two key observations to refine the calculation of P(S | T):

  1. Preference Order: This factor models the technician's diagnostic and resolution process. Technicians, whether novice or expert, follow a certain sequence when accessing configuration blocks. An expert technician can typically identify and access the root cause block more directly, placing it higher in their mental "preference order" of investigation. Heimdall abstracts this as a preference order defined on the configuration blocks within the RDG subgraph relevant to the ticket. This allows the model to differentiate risk based on the assumed skill level and diagnostic path of the operator.
  1. Root Cause Estimation: Before a ticket is truly resolved, the exact root cause is often an educated guess. Heimdall's model incorporates root cause estimation, which is computed using historical statistics from past tickets and resolutions. This predictive element allows the system to anticipate potential high-risk areas and prioritize monitoring based on the most likely source of the problem. While historical data is the primary input, the framework is extensible to incorporate other estimation methods, such as those from advanced debugging tools.

These three elements—the RDG, preference order, and root cause estimation—collectively contribute to the accurate determination of P(S | T), the detailed decomposition of which is further elaborated in the associated paper.

Risk Monitoring and Response System

The final component is the Risk Monitoring and Response System, which acts as a centralized reference monitor. Its primary function is to inspect every command issued by a third-party technician, identify the affected configuration blocks, and make a real-time decision on whether the command should be permitted.

The decision-making process is guided by a Risk Response Policy (RRP), defined by the enterprise, in conjunction with the current accumulated risk. When access to modify a block B is requested, the system calculates the current risk by accumulating all risks associated with blocks B' that are placed earlier in the technician's preference order than B. This cumulative risk reflects the potential impact of the diagnostic path taken so far.

The system then refers to the RRP to check if the current risk exceeds a predefined threshold. If it does, corresponding actions are triggered, such as alerting the enterprise administrator or immediately stopping the command execution. To streamline the workflow, the RRP can also define access granting badges, which allow enterprise administrators to pre-approve certain levels of access based on ticket type or technician profile, expediting legitimate troubleshooting. This dynamic, real-time enforcement mechanism ensures that network changes align with the enterprise's risk tolerance.

Demo / Proof of Concept

▶ Watch: Observation 1: Operator access order affects consequence likelihood (6:20)

While the talk did not feature a live, interactive demonstration of Heimdall in action, the speaker presented a comprehensive evaluation and validation of the system's capabilities and performance. This extensive study effectively serves as a proof of concept, demonstrating Heimdall's practical applicability and efficacy in real-world scenarios.

The evaluation involved a comprehensive expert validation consisting of 99 rounds of experiments conducted over 41 hours. These experiments aimed to demonstrate Heimdall's efficiency for practical usage and its ability to reduce outsourcing risks effectively.

Key results from this validation included:

  • Risk Reduction Efficacy: Heimdall demonstrated significant effectiveness in mitigating outsourcing risks. A striking 92% of tasks were solved without incurring any extra risks beyond those directly associated with the identified root cause block. This highlights the system's precision in isolating and containing the impact of troubleshooting activities.
  • Accuracy of Risk Assessment: The evaluation specifically compared Heimdall's RDG construction method against traditional approaches. It was found that purely token-matching methods generate significantly more false positive dependency links than Heimdall's data plane-leveraged approach, affirming the superior accuracy of Heimdall's risk assessment model.
  • Performance of System Components:
  • Reference Monitor: The centralized reference monitor proved highly efficient, processing commands with a negligible overhead of only four to five milliseconds per command. This low latency is crucial for maintaining a smooth privilege management workflow.
  • Workflow Overhead: Overall, 86% of tasks incurred less than 10% overhead on the risk-aware privilege granting workflow, indicating that Heimdall can be integrated without significantly slowing down technician operations.
  • Scalability: The RDG construction and subsequent risk computations were shown to scale easily to hundreds of routers, making Heimdall suitable for large backbone networks. Importantly, all such computations are performed proactively before a technician begins work, ensuring that the system does not introduce delays during active troubleshooting.

The comprehensive nature of these experiments, including studies on affecting factors of risks such as block access order and privilege granting granularity (further detailed in the paper), collectively validates Heimdall as a robust and practical solution for risk-aware network management outsourcing.

Defensive Implications

▶ Watch: Observation 2: Root cause estimation using historical data (7:40)

Heimdall presents several crucial defensive implications for enterprises engaging in or considering outsourced network management:

  1. Adopt Risk-Aware Outsourcing Frameworks: Enterprises should move beyond traditional, static access control models and integrate dynamic, risk-aware frameworks like Heimdall. This means evaluating third-party access not just on "what" commands can be run, but "where" and "when" they are run, and "what impact" they could have on specific assets.
  1. Implement Granular, Context-Sensitive Access Control: The coarse-grained privilege levels offered by many network devices are insufficient. Defenders should advocate for and implement systems that provide granular access control based on the specific context of a troubleshooting ticket, the technician's likely path, and the real-time accumulated risk. This allows for a "least privilege" model that adapts dynamically.
  1. Leverage Data Plane Information for Dependency Mapping: Relying solely on configuration file parsing for understanding network dependencies is prone to errors and false positives. Security teams should explore tools and methodologies that incorporate data plane information to build more accurate risk dependency graphs. This helps in precisely identifying the blast radius of any configuration change.
  1. Incorporate Behavioral Analytics and Historical Data: The talk highlights the importance of technician preference order and root cause estimation from historical data. Defenders should collect and analyze historical troubleshooting data to identify common problem patterns, assess technician expertise, and refine risk models. This allows for proactive risk mitigation based on predictive insights.
  1. Establish Clear Risk Response Policies (RRPs): Enterprises must define clear and actionable Risk Response Policies that dictate automated responses when accumulated risk thresholds are exceeded. These policies should include alerting mechanisms, automatic command blocking, and escalation procedures to ensure rapid containment of potential incidents.
  1. Prioritize Proactive Risk Computation: The fact that Heimdall performs RDG construction and risk computations proactively, before technician engagement, is a significant advantage. Defenders should seek solutions that minimize operational overhead during critical incident response by pre-calculating potential risks and dependencies.
  1. Consider "Access Granting Badges" for Efficiency: To balance security with operational efficiency, enterprises can implement mechanisms similar to Heimdall's access granting badges. This allows for predefined, risk-assessed access profiles for common tasks, streamlining the approval process for trusted operations while still maintaining risk oversight.

By integrating these defensive strategies, organizations can significantly reduce their exposure to risks associated with outsourced network management, transforming a potential vulnerability into a more secure and controlled operational model.

Key Takeaways

  • Traditional network risk models, which associate risk with individual commands, are inflexible and coarse-grained, making them inadequate for securing outsourced network management.
  • Heimdall introduces an asset-based quantitative risk model that defines risk based on the potential impact on valuable enterprise assets, leveraging conditional probabilities and asset values.
  • The framework constructs highly accurate Risk Dependency Graphs (RDGs) by incorporating data plane information, significantly reducing false positive dependencies compared to token-matching approaches.
  • Heimdall enhances risk assessment by modeling technician behavior through preference orders and utilizing historical data for root cause estimation, providing a more nuanced understanding of potential risks during troubleshooting.
  • A centralized reference monitor enforces risk response policies in real-time, inspecting every command, accumulating risk based on the technician's actions, and performing actions like alerting or blocking when thresholds are exceeded.
  • Evaluations demonstrate Heimdall's efficiency (4-5ms/command overhead, 86% tasks with <10% overhead) and effectiveness (92% of tasks solved without extra risks), proving its scalability for large networks and practical applicability.

About the Speaker(s)

The talk "Heimdall: Towards Risk-Aware Network Management Outsourcing" was presented by Yuejie Wang. No further specific biographical details, such as their title or company affiliation, were provided within the scope of the presentation transcript.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Heimdall is legitimate academic systems work solving a real problem — third-party privileged access to production networks is genuinely underserved by existing tooling. The data-plane-leveraged RDG construction is the most interesting piece, and the quantitative asset-based risk model is a step up from flat command-privilege mappings. But this is a conference paper presentation, not a practitioner talk, and it shows: the depth stays at the model description level, the demo is entirely absent, and the results feel clean in a way that real messy production networks rarely are.

Heather Calloway (CISO) — SOLID

Heimdall addresses a real and underexamined risk — privileged third-party access during network troubleshooting — with a technically credible, quantitative framework. The research is sound and the problem is institutionally significant, but the talk stays in the research register and never crosses into the governance or operational decisions that would make it actionable for the CISOs and risk owners who most need it.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025