SketchFeature: High-Quality Per-Flow Feature Extractor Towards Security-Aware Data Plane
Sian Kim (Ewa Women's University)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Internet Security
Overview
In an era where network security increasingly relies on sophisticated AI-enhanced in-network defense mechanisms, the ability to efficiently extract high-quality packet features has become paramount. Yang Kim from Ewa Women's University presented "SketchFeature: High-Quality Per-Flow Feature Extractor Towards Security-Aware Data Plane" at the NDSS Symposium, addressing critical shortcomings in current network monitoring capabilities. The talk highlights that incomplete or low-resolution feature extraction can lead to significant blind spots and misclassifications, leaving networks vulnerable to a wide array of attacks. For rapid attack mitigation, real-time feature extraction directly within the data plane is indispensable, yet it remains a formidable challenge due to the inherent hardware limitations of network devices.
Key moments
- 0:00 Introduction: The need for high-quality packet features
- 2:50 Existing solutions' limitations and ideal feature extraction requirements
- 3:20 Introducing SketchFeature: High-resolution, All-flow, Full-range
- 4:00 Overview of SketchFeature's theoretical and experimental validation
- 6:20 Key innovation: Sketch virtualization for efficient memory utilization
- 8:00 Key innovation: Membership test preventing phantom decoding
SketchFeature: High-Quality Per-Flow Feature Extractor Towards Security-Aware Data Plane
Speakers: Sian Kim (Ewa Women's University)
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=rWfltIN9Y_I
Overview
In an era where network security increasingly relies on sophisticated AI-enhanced in-network defense mechanisms, the ability to efficiently extract high-quality packet features has become paramount. Yang Kim from Ewa Women's University presented "SketchFeature: High-Quality Per-Flow Feature Extractor Towards Security-Aware Data Plane" at the NDSS Symposium, addressing critical shortcomings in current network monitoring capabilities. The talk highlights that incomplete or low-resolution feature extraction can lead to significant blind spots and misclassifications, leaving networks vulnerable to a wide array of attacks. For rapid attack mitigation, real-time feature extraction directly within the data plane is indispensable, yet it remains a formidable challenge due to the inherent hardware limitations of network devices.
This presentation introduces SketchFeature, a novel scheme designed to overcome these limitations by enabling high resolution, all flow, and full range (half) feature measurement. Unlike previous approaches that either selectively monitor flows or store features at low granularity, SketchFeature aims to extract comprehensive per-flow distribution features with superior accuracy and efficiency. By introducing innovative strategies such as sketch virtualization and a membership test, SketchFeature promises to enhance attack detection capabilities, reduce false positives, and ensure robust, line-rate processing even under demanding network conditions. Its demonstrated deployability in commercial switches, like the Tofino, underscores its practical relevance for advancing network security.
Background
▶ Watch: Introduction: The need for high-quality packet features (0:00)
The landscape of modern network security is increasingly shaped by the integration of artificial intelligence and machine learning. These advanced defense systems rely heavily on the quality and timeliness of the network traffic features they analyze. Features can be broadly categorized by their dimensionality. First order features are per-packet attributes, such as protocol type and Time-To-Live (TTL). Second order features capture per-flow status, including flow size, mean packet length, and variance. Crucially, third order features provide deeper insights by capturing per-flow distributions, such as inter-packet delay distribution or packet size distribution. It is widely acknowledged that these higher-order features significantly enhance the capabilities of attack detection algorithms, allowing for the identification of more subtle and sophisticated threats.
However, extracting these rich, higher-order features directly within the data plane presents substantial challenges. The data plane, responsible for forwarding packets at wire speed, operates under stringent hardware limitations. It must guarantee line rate packet processing, which restricts computational operations to simple arithmetic and imposes severe constraints on memory resources, thereby preventing the storage of detailed per-flow distribution features. To circumvent these limitations, existing approaches often employ a symptom detector mechanism. These detectors selectively extract features from a subset of flows, escalating only those exhibiting predefined "symptoms" to the control plane for more in-depth analysis and attack detection.
Among such symptom detectors, Flow Lens and Net Warden are two notable prior works. Flow Lens attempts to capture per-flow distributions by using a hash table to quantize feature ranges and storing only the top K most frequent bins. While an improvement, this method is susceptible to evasion; attackers can subtly shift packet size patterns to avoid falling into the top K bins, rendering the system unable to generalize to unseen attack variants. This "top K bias" limits its effectiveness against adaptive adversaries.
Net Warden, on the other hand, utilizes a sketch-based approach to store per-flow distributions. While more comprehensive than Flow Lens in its intent, Net Warden operates at a relatively low resolution. This low-resolution quantization leads to two significant problems. First, small variations in feature values might fall into the same bin, resulting in misclassification and reduced detection accuracy. Second, its low resolution makes it vulnerable to control plane flooding attacks, where adversaries can intentionally trigger an excessive number of flow escalations to overwhelm the more resource-intensive control plane, disrupting normal operations and potentially creating a denial-of-service condition.
Both Flow Lens and Net Warden ultimately suffer from the fundamental constraints of the data plane, failing to provide the comprehensive, high-resolution feature extraction necessary for robust AI-enhanced network defense. The imperative, therefore, is to develop a better approach that moves beyond selective feature extraction from a small subset of flows, covering all flows, storing the full feature distribution range instead of just top K bins, and achieving high resolution while minimizing memory and computational overhead to sustain line-rate processing. This is precisely the gap that SketchFeature aims to fill.
Key Findings
▶ Watch: Introducing SketchFeature: High-resolution, All-flow, Full-range (3:20)
SketchFeature represents a significant advancement in data plane feature extraction by delivering on the promise of high resolution, all flow, and full range (half) feature measurement. This novel approach fundamentally addresses the limitations of prior work, eliminating the "top K bias" and enabling the detection of a wider spectrum of attacks with unprecedented granularity and accuracy directly within network hardware.
The core of SketchFeature's innovation lies in two distinct yet complementary strategies: sketch virtualization and a novel membership test. Sketch virtualization efficiently manages memory utilization and load balancing for feature storage, even under highly skewed traffic distributions, a common challenge in real-world networks. The membership test, on the other hand, rigorously resolves the critical issue of "phantom decoding," where non-existent feature values might be erroneously reported due to hash collisions, thereby significantly reducing false positives and enhancing overall accuracy.
The efficacy of SketchFeature is not merely theoretical; its error bounds for the encoding and decoding processes have been rigorously proven, and its accuracy and efficiency have been extensively validated through experimental deployments. Comparative analyses against existing schemes across various security use cases—including cover channel, DDoS, and botnet detection—have consistently demonstrated SketchFeature's superior performance in producing feature distributions that closely approximate perfect measurements. A pivotal finding is its successful deployment and operation in a Tofino commercial switch, unequivocally showcasing its practical applicability and readiness for real-world network security infrastructure.
Technical Deep Dive
▶ Watch: Overview of SketchFeature's theoretical and experimental validation (4:00)
SketchFeature's design meticulously addresses the challenges of data plane limitations by combining efficient data structures with innovative processing strategies. The fundamental goal is to represent complex per-flow distributions in a manner suitable for high-speed, memory-constrained hardware.
The process begins with quantization, a crucial step where continuous feature values (e.g., inter-packet delays, packet sizes) are transformed into discrete bins. This conversion is essential to cope with the limited memory resources of ASIC-based switches. For each flow, the system records how many packets fall into each bin, creating a distribution vector. This vector provides a binned representation of the flow's characteristics.
To store and retrieve these distribution vectors efficiently for all flows, SketchFeature employs a specialized data structure comprising a sketch and a Bloom filter.
The encoding process works as follows: when a packet arrives, its relevant feature value (e.g., packet size) is first quantized to determine its corresponding bin number. This bin number serves as an input to multiple hash functions. These hash functions are then used to interact with both the Bloom filter and the sketch. For the Bloom filter, the bits corresponding to the hash outputs are set to one, indicating the presence of a feature in that specific bin for the flow. Simultaneously, for the sketch, the counters at the hashed locations are incremented by one, recording the occurrence of the feature.
The decoding process is initiated when a query for a specific flow ID and bin number is made. First, the Bloom filter is consulted to verify whether the feature is actually present in the given bin. If the Bloom filter indicates presence, the value is then retrieved from the sketch. SketchFeature utilizes a count-mean sketch, a probabilistic data structure that estimates frequencies. To minimize errors inherent in sketch operations, it obtains the minimum value across all layers of the sketch. After decoding all relevant bins, the full distribution vector for the queried flow can be reconstructed.
A core innovation embedded within this design is sketch virtualization. Traditional approaches, termed sketch partitioning, allocate a separate, physically isolated sketch for each discrete bin. While conceptually simple, this method is highly inefficient and prone to accuracy degradation. If a specific bin contains a disproportionately large amount of data (a common scenario under skewed feature distributions, such as during an attack), its dedicated sketch can become heavily congested, leading to significant inaccuracies. Sketch virtualization, in contrast, shares a single, large sketch across all quantized bins. It distinguishes between bins using robust hash functions, effectively creating a virtualized space within the physical sketch. This approach inherently provides a form of load balancing; due to the uniform characteristics of hash functions, data is distributed more evenly across the shared sketch, regardless of the underlying feature distribution. This leads to more efficient memory utilization and, crucially, higher accuracy, especially under skewed traffic conditions where sketch partitioning fails dramatically. Experimental results vividly illustrate this, showing sketch partitioning with heavily populated, congested bins (indicated in red), while sketch virtualization maintains a much more even distribution, ensuring accurate results.
The second pivotal innovation is the membership test, designed to resolve the problem of phantom decoding. Phantom decoding occurs when a query retrieves a non-existent bin value due to hash collisions within the sketch. In essence, even if a particular bin was never actually encoded with a value, the sketch might return a non-zero count because other, unrelated data happened to hash to the same location. This leads to false positives, where the system believes a feature exists when it does not. SketchFeature tackles this by applying the membership test before querying the sketch. By first checking the Bloom filter, it ensures that the queried bin actually exists and has been encoded at some point. Only if the Bloom filter confirms presence does the system proceed to retrieve the value from the sketch. This pre-check significantly reduces the false positive rate, dramatically improving the overall accuracy of feature extraction. Visualizations presented in the talk clearly demonstrate that without the membership test, false positive errors (yellow) can overwhelm true positive errors (gray), but with its application, false positives are effectively reduced, enhancing the fidelity of the extracted features.
The quality of the per-flow distributions extracted by SketchFeature was quantitatively evaluated using the Weighted Relative Error (WRE) metric. A WRE value closer to zero indicates higher accuracy. The results consistently showed that applying sketch virtualization alone significantly enhances measurement accuracy, and incorporating the membership test further refines it, leading to the lowest WRE values. When comparing the traffic distribution quality of SketchFeature against a baseline and the actual distribution, the baseline exhibited errors exceeding a factor of 10 in the worst case (on a logarithmic scale). In stark contrast, SketchFeature's output closely aligned with the actual distribution, showcasing its significantly improved accuracy and reliability for security applications.
Demo / Proof of Concept
▶ Watch: Key innovation: Sketch virtualization for efficient memory utilization (6:20)
The practical applicability and superior performance of SketchFeature were rigorously demonstrated through a series of comprehensive evaluations and a real-world deployment. The most compelling proof of concept involved its successful implementation in a Tofino commercial switch, a high-performance programmable network device. This deployment confirms that SketchFeature's intricate design, including sketch virtualization and the membership test, can operate effectively at line rate within constrained data plane environments.
The evaluation focused on assessing the quality of extracted features, specifically inter-packet delay (IPD) and packet size distribution (PCI), under various datasets and with different memory allocations. Across all tested memory sizes and workloads, SketchFeature consistently and significantly outperformed the baseline methods in terms of feature measurement accuracy, as quantified by the Weighted Relative Error (WRE) metric. This robustness across diverse operational parameters highlights its reliability in varied network conditions.
Beyond raw feature quality, the ultimate test of SketchFeature's utility lies in its ability to enhance attack detection. The research team conducted extensive experiments across multiple critical security use cases: cover channel detection, DDoS (specifically using the CI/CDOS data set for attack traffic and the Kada data set for benign traffic), and botnet detection. For these evaluations, features were extracted using four different schemes—NetBeacon, NetWarden, Flow Lens, and SketchFeature—alongside a "perfect measurement" baseline (representing the actual, uncompressed feature distribution). These extracted features were then fed as input into a CNN model, which was trained to classify attack traffic.
The results unequivocally demonstrated SketchFeature's superior detection performance. The feature distributions produced by SketchFeature were consistently closest to the perfect measurement, leading to the highest AU accuracy and F1 scores, coupled with the lowest false positive rate (FPR) and false negative rate (FNR) across all tested attack types. For instance, in DDoS detection, by providing more granular and accurate inter-packet delay and packet size distributions, SketchFeature enabled the CNN model to distinguish malicious traffic patterns from benign ones with greater precision than any other evaluated method. This comprehensive validation underscores SketchFeature's capability to provide superior feature extraction, directly translating into enhanced security for network defense systems.
Defensive Implications
▶ Watch: Key innovation: Membership test preventing phantom decoding (8:00)
SketchFeature offers profound implications for network defenders seeking to bolster their security posture against increasingly sophisticated threats. Its ability to provide high-quality, real-time per-flow feature extraction directly in the data plane presents several actionable strategies:
- Upgrade Feature Extraction Capabilities: Organizations should prioritize the deployment of advanced feature extraction mechanisms in their network data plane. Implementing solutions like SketchFeature, which ensures high resolution, all flow, and full range (half) feature measurement, moves beyond traditional, low-granularity monitoring. This provides a richer, more accurate dataset for downstream security analytics.
- Leverage Higher-Order Features for Enhanced Detection: Defenders can now reliably extract and utilize third order features, such as inter-packet delay and packet size distributions, which are significantly more effective for detecting subtle and advanced attacks (e.g., cover channels, sophisticated botnet C2 traffic) compared to simpler first and second-order features. This allows security operations centers (SOCs) to develop more nuanced detection rules and machine learning models.
- Mitigate Evasion and Control Plane Flooding: By eliminating the "top K bias" inherent in methods like Flow Lens, SketchFeature makes it significantly harder for attackers to evade detection by slightly altering their traffic patterns. Furthermore, its high-resolution data plane processing reduces the reliance on escalating numerous flows to the control plane, thereby mitigating the risk of control plane flooding attacks that Net Warden is susceptible to. This ensures the control plane remains responsive for critical security decisions.
- Harness Programmable Switch Capabilities: The successful deployment of SketchFeature on a Tofino commercial switch highlights the critical role of programmable network hardware. Defenders should explore and invest in programmable switches to implement advanced, custom security functions directly at line rate, enabling proactive and real-time threat response without impacting network performance.
- Enhance AI/ML-Driven Security Analytics: The accurate and comprehensive feature sets provided by SketchFeature serve as superior input for AI/ML models used in security analytics. Cleaner, more detailed data leads to more robust model training, improved classification accuracy, and a significant reduction in both false positive rates (FPR) (reducing alert fatigue) and false negative rates (FNR) (preventing missed threats). This directly translates to more effective and efficient security operations.
In essence, SketchFeature empowers defenders to move from reactive, coarse-grained security to a proactive, granular, and AI-optimized defense strategy, making networks more resilient to evolving cyber threats.
Key Takeaways
- AI-enhanced network defense critically depends on high-quality, real-time feature extraction directly within the data plane, a challenge due to hardware limitations.
- Existing symptom detectors like Flow Lens and Net Warden suffer from significant drawbacks, including top K bias, low resolution, vulnerability to evasion, and susceptibility to control plane flooding attacks.
- SketchFeature introduces "half" (high resolution, all flow, full range) feature measurement, overcoming prior limitations to provide comprehensive and accurate per-flow distribution data.
- Key innovations are sketch virtualization, which ensures efficient memory utilization and load balancing under skewed traffic, and a membership test, which resolves phantom decoding to significantly reduce false positives.
- Experimental validation, including deployment on a Tofino commercial switch, confirms SketchFeature's superior accuracy (lower WRE) and enhanced attack detection performance (AU accuracy, F1 score, FPR, FNR) across various security use cases (cover channel, DDoS, botnet detection).
- Defenders can leverage SketchFeature to upgrade feature extraction, utilize higher-order features, mitigate evasion and flooding, and provide superior input for AI/ML security models, leading to more robust and efficient network defense.
About the Speaker(s)
The paper "SketchFeature: High-Quality Per-Flow Feature Extractor Towards Security-Aware Data Plane" was presented by Yang Kim from Ewa Women's University.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate systems research with a concrete contribution — sketch virtualization plus a Bloom-filter membership test to suppress phantom decoding is a real, implementable idea with provable error bounds and Tofino validation. It's solid conference-grade work, but it's incremental: the building blocks (count-min sketch, Bloom filters, programmable ASICs) are all well-established, and the threat models addressed (evasion of top-K detectors, control plane flooding) are known problems getting an engineering fix rather than a conceptual breakthrough.
Heather Calloway (CISO) — WEAK
Technically sound systems research on data plane feature extraction, with real engineering merit and a working Tofino deployment. But this talk stops at the lab bench — it never crosses into institutional relevance, operational decision-making, or anything a security leader needs to act on.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025