Advancing Network Threat Detection Thru Standardized Feature Extraction & Dynamic Ensemble Learning
Jason Ford (Principal Research Engineer · Proofpoint)
BSides Las Vegas 2025 · Day 1
Overview
Jason Ford, introducing himself as a research engineer at Proofpoint giving his first BSides talk, presents roughly two years of research on improving network intrusion detection by fixing what he argues is the real bottleneck in many machine-learning NIDS efforts: feature extraction and generalization. The talk contrasts classical signature-driven NIDS with NDR-style behavioral analytics, then proposes a three-part pipeline: a standardized feature framework derived from raw pcaps, training a basket of diverse classifiers, and combining them with a custom ensemble method he names after himself—“Ford class” (expanded in-slide as Ford class-specific weighted values). He reports validation accuracy just under 98% for the ensemble, with emphasis on balanced precision and recall relative to individual models, and outlines future work including hyperparameter tuning for weaker models and field testing using a Raspberry Pi with channel-hopping Wi-Fi capture in a conference hotel environment.

Key moments
- 2:00 Speaker intro: Proofpoint research engineer, first BSides talk; problem statement—ML NIDS often fails to generalize beyond small lab datasets.
- 4:00 NIDS vs NDR definitions: signatures/reputation vs behavioral analytics, anomaly models, hunting/response (Vectra mentioned).
- 6:00 Core thesis: feature extraction often matters more than model choice; IP/port features and auto feature learning both have limits on temporal traffic.
- 8:00 Data sourcing: public PCAP datasets from three universities plus home/academic benign captures to balance training.
- 10:00 Feature families: TCP flag/TTL-style packet stats plus flow forward/back timing and windowed counts aimed at generalization.
- 12:00 Model zoo + ensemble plan; Isolation Forest’s role when benign is well-known; weaker GMM/NN flagged for hyperparameter work.
- 14:00 ‘Ford class’ ensemble math: weight models by validation accuracy; up-weight uncertain (near 0.5) predictions vs overconfident extremes.
- 16:00 Results table: ensemble ~97–98% accuracy with balanced precision/recall; Pi-based live-capture experiment photo teased for next steps.
Advancing Network Threat Detection Thru Standardized Feature Extraction & Dynamic Ensemble Learning
Speakers: Jason Ford, Research Engineer, Proofpoint
Conference: BSides Las Vegas
YouTube: https://www.youtube.com/watch?v=_06VmadFSTI
Overview
Jason Ford, introducing himself as a research engineer at Proofpoint giving his first BSides talk, presents roughly two years of research on improving network intrusion detection by fixing what he argues is the real bottleneck in many machine-learning NIDS efforts: feature extraction and generalization. The talk contrasts classical signature-driven NIDS with NDR-style behavioral analytics, then proposes a three-part pipeline: a standardized feature framework derived from raw pcaps, training a basket of diverse classifiers, and combining them with a custom ensemble method he names after himself—“Ford class” (expanded in-slide as Ford class-specific weighted values). He reports validation accuracy just under 98% for the ensemble, with emphasis on balanced precision and recall relative to individual models, and outlines future work including hyperparameter tuning for weaker models and field testing using a Raspberry Pi with channel-hopping Wi-Fi capture in a conference hotel environment.
Background
▶ Watch: Speaker intro: Proofpoint research engineer, first BSides talk; problem state... (2:00)
Ford frames the problem as adversary-driven complexity: attacks are harder to detect, while traditional NIDS remains heavily influenced by reputation and signatures—strong on known bad, weaker on novelty. He cites prior ML-in-NIDS literature as often training single classifiers on small datasets that look excellent in papers but fail in the wild, sometimes reporting under 40% efficacy when exposed to broader traffic.
He briefly defines NIDS versus NDR for the room: NIDS as more rule/signature oriented; NDR as incorporating behavioral analytics, commonly anomaly detection, plus capabilities such as threat hunting and automated response in commercial products (he names Vectra as an example). His research comparison target remains NIDS, but the feature and ensemble ideas are portable to broader detection stacks.
To understand why this distinction matters for enterprises, consider what each system optimizes for under stress. Signature systems optimize for precision on known bad and explainability—analysts like matches that map to a name, a campaign, a CVE. Behavioral systems optimize for coverage of unknowns—anomalies that might be benign operational weirdness or the first heartbeat of something new. Ford’s work sits in the uncomfortable middle ground that many SOCs actually need: statistically learned judgments on the wire, but grounded in features that do not instantly expire when NAT pools rotate or when cloud egress IPs change hourly.
The talk also quietly engages a cultural problem in applied ML security: the gap between benchmark culture and operations culture. Benchmark culture rewards a single number and a leaderboard. Operations culture rewards stability under drift, low false-positive burden, and the ability to escalate with evidence that an on-call analyst can sanity-check at 3 a.m. Ford’s emphasis on precision/recall balance and skepticism about “one classifier to rule them all” is aligned with operations—even if academic incentives still push pretty ablations on toy CSVs.
Key Findings
▶ Watch: Core thesis: feature extraction often matters more than model choice; IP/port... (6:00)
Features dominate models (in practice). Ford argues that in many ML security projects, poor generalization traces back to bad features, not only model choice. Prior work over-relied on volatile or overly specific fields like IP addresses and port numbers, which do not survive real-network diversity. Another common approach—letting a model learn features end-to-end—can miss temporal structure needed to separate benign from malicious transactional traffic.
Raw pcaps beat stripped CSVs for research flexibility. Public datasets often publish reduced feature CSVs without the full packet context. Ford sourced datasets that include raw pcaps so he could compute flow statistics, directional packet behavior, timing, and richer metadata.
Dataset composition required balancing benign reality. He names three academic PCAP sources: Czech Technical University, University of New South Wales, and University of Science and Technology of China, plus additional benign captures from home and academic networks to represent ordinary applications (video calls, chat tools, web browsing) that attack-heavy corpora omit.
Feature families emphasize generalization. At packet granularity he mentions engineered quantities like TCP flag counts and average TTL. At flow level he emphasizes forward/backward statistics, inter-arrival times, counts of packets observed in recent windows (he references a 10-second window in Q&A), duration, and related flow metadata—intentionally steering away from features that collapse when IPs/ports rotate.
Classifier zoo + ensemble beats single-model fragility. He trains multiple models including Random Forest, Isolation Forest (common in NDR-ish anomaly setups), Gaussian Mixture Model (GMM), Quadratic Discriminant Analysis (QDA), boosting methods, and neural networks. Some weaker performers (GMM, NN) are flagged for future hyperparameter work; CNN/RNN promise from other studies is noted as future exploration.
Novel ensemble weighting targets overconfidence. The ensemble combines per-model benign/malicious scores using validation accuracies as base weights, with a twist: predictions near 0.5 receive higher weight than predictions near 0 or 1, explicitly to down-rank overconfident classifiers. Final ensemble score is described as a weighted average compared to a 0.5 decision threshold.
Reported results: ensemble accuracy ~97–98% with the most balanced precision/recall among tested approaches in his evaluation table (individual boosting models also did well). Latency: in Q&A he claims classification decisions under 30 milliseconds on a Raspberry Pi 4 using a sliding window framing (described as 10 seconds before/after around a point—total 30 seconds of traffic context in the answer, wording slightly ambiguous in the recording).
Another finding implicit in the narrative is that model families differ in failure modes. Isolation Forest is attractive when you have a strong model of benign and an incomplete malicious corpus—common in real enterprises. Boosted ensembles often perform well on tabular features but can overfit if labels are noisy or if training data accidentally includes proxy leakage. Neural networks can capture nonlinear interactions but, as Ford notes, become hyperparameter-sensitive and can underperform if the feature tensor does not represent temporal structure well. The ensemble story is less “collect Pokémon models” and more “combine uncorrelated errors.”
Finally, Ford is candid that encrypted traffic changes what you can responsibly claim. The hotel Wi-Fi experiment explicitly avoids inspecting encrypted payloads; a trunk-port deployment would still face TLS ubiquity on the internet. The defensive point remains: metadata and flow dynamics still leak behavioral structure, but anyone selling “ML on the wire solves everything” without stating encryption constraints is selling snake oil. Ford’s feature focus on timing, directionality, and TCP-level statistics is consistent with a metadata-first posture—though he does not overpromise decryption.
Technical Deep Dive
▶ Watch: Feature families: TCP flag/TTL-style packet stats plus flow forward/back timi... (10:00)
Ford walks the pipeline as boxes: load samples → extract features → scale → per-model predictions → aggregate ensemble score → classify.
The feature extraction framework is positioned as the reusable artifact. By operating from pcaps, the same code paths can reconstruct flows and directional series rather than trusting precomputed feature dumps that may hide label leakage or selection bias.
The ensemble section is the mathematical heart. After validation, each classifier gets a weight derived from performance, then predictions are fused. The anti-overconfidence weighting is presented as empirically strong—implicitly arguing many false positives/negatives in security ML come from models that are confidently wrong on out-of-distribution traffic.
Scaling also deserves emphasis because network flow features can span orders of magnitude: durations, packet counts, and inter-arrival times do not naturally live on comparable scales. Without normalization, some models will effectively “listen” only to the largest-magnitude features, creating brittle decisions. Ford mentions applying scaling after extraction; while he does not deep-dive sklearn specifics on stage, the engineering intent is standard good practice—make the learning problem stable.
The Q&A exchange about retraining and dynamic weights clarifies intent versus current artifact: today’s ensemble is static relative to training data, but a production-minded roadmap would add continuous or periodic retraining, careful drift monitoring, and governance around label quality. That is how an interesting research prototype becomes something a SOC manager will tolerate.
He is careful to separate lab evaluation from production truth: the live Pi experiment is portrayed as an early feasibility probe, sniffing open Wi-Fi (with a wry note about observing casino traffic during hacker week), not a finished product detector.
Demo / Proof of Concept
▶ Watch: Model zoo + ensemble plan; Isolation Forest’s role when benign is well-known;... (12:00)
No live demo is shown in the recorded talk segment reviewed; instead, he displays result tables and a photo of a Raspberry Pi with an external Wi-Fi adapter for channel hopping on open networks. Q&A clarifies intentions: trunk port placement would be ideal for enterprise deployment rather than Wi-Fi-only capture.
Defensive Implications
▶ Watch: Results table: ensemble ~97–98% accuracy with balanced precision/recall; Pi-b... (16:00)
For blue teams, the actionable thesis is to scrutinize ML claims through feature generalizability and operational windows, not headline accuracy on static corpora. If your model’s features are essentially addresses and ports, it will age poorly. If your model cannot explain why a flow scored malicious, you still need analyst workflows—Ford notes that even with a score, determining “what it is” may be a second-phase problem.
For researchers, he released the feature framework on GitHub under GPL (URL shown on slides; not transcribed verbatim here). That licensing choice has downstream implications for commercial integration—teams must respect GPL obligations if they fork or embed.
For detection engineering, the ensemble’s precision/recall balance is the right optimization target for noisy networks: accuracy alone is a lie when base rates skew benign.
Key Takeaways
- Signature-heavy NIDS remains necessary but insufficient; ML help must generalize beyond lab sets.
- Feature engineering from raw pcaps and flow metadata is framed as more stable than address/port centric training.
- Balanced benign data matters: attack-only corpora teach models the wrong world distribution.
- Ensemble methods can reduce single-model weaknesses; Ford’s weighting scheme explicitly punishes overconfident extremes.
- Reported offline performance reaches ~97–98% accuracy with strong precision/recall balance in the speaker’s tests.
- Edge deployment may be feasible: sub-30ms inference is claimed on a Pi 4 in Q&A, pending real-world validation.
- Future work includes better NN/GMM tuning, possible CNN/RNN extensions, and meaningful evaluation on live traffic—not only labeled pcaps.
About the Speaker(s)
Jason Ford is a research engineer at Proofpoint with roughly 10+ years at the company, per his introduction. He mentions prior published work including ML for QR code detection in email and research involvement on state-sponsored actor problems. He cites education at the University of South Carolina and credits a mentor (Sherri DeGrio at Microsoft threat research, spelling as spoken) for pushing him to name a favorite protocol—he chooses DNS as foundational and failure-prone. Personal details include living in Colorado and enjoying outdoor activities.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
A grounded ML-for-NIDS research talk with the right skepticism about dataset realism, a reproducible-ish pipeline story, and an ensemble trick that is simple enough to implement but not trivial to invent—exactly the BSides bar for applied research.
Heather Calloway (CISO) — STRONG ACCEPT
This is the kind of research narrative procurement and SOC leadership should demand before buying ‘ML NIDS’ theater: emphasize generalizable features, report precision/recall balance, and separate offline accuracy from production drift and analyst workload.