Endangered Privacy: Large-Scale Monitoring of Video Streaming Services
Martin Björklund
34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Privacy 3: Attacks
Overview
In a revealing presentation at USENIX Security, Martin Björklund unveiled groundbreaking research demonstrating that a sophisticated man-in-the-middle (MitM) eavesdropper can precisely identify the specific video content a user is watching on major streaming services, even though the stream itself is encrypted. Titled "Endangered Privacy: Large-Scale Monitoring of Video Streaming Services," the talk showcased how inherent properties of modern video streaming protocols, specifically Adaptive Bit Rate (ABR) streaming and Variable Bit Rate (VBR) encoding, create unique "fingerprints" from network traffic that can be exploited for large-scale surveillance.

Key moments
- 0:00 Introduction: Encrypted video privacy vulnerability
- 1:30 Variable bitrate encoding creates unique video fingerprints
- 2:15 Bursty network traffic reveals exact video content
- 3:45 Project goals: Practical, large-scale, protocol-agnostic attack
- 5:45 Efficient data collection via manifest files
- 6:20 KD-tree for fast fingerprint matching
- 7:00 Live attack: T-shark and heuristic for identification
- 8:40 Real-time demonstration of video identification confidence
Endangered Privacy: Large-Scale Monitoring of Video Streaming Services
Speakers: Martin Björklund
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=x-8U__f9Z6E
Overview
In a revealing presentation at USENIX Security, Martin Björklund unveiled groundbreaking research demonstrating that a sophisticated man-in-the-middle (MitM) eavesdropper can precisely identify the specific video content a user is watching on major streaming services, even though the stream itself is encrypted. Titled "Endangered Privacy: Large-Scale Monitoring of Video Streaming Services," the talk showcased how inherent properties of modern video streaming protocols, specifically Adaptive Bit Rate (ABR) streaming and Variable Bit Rate (VBR) encoding, create unique "fingerprints" from network traffic that can be exploited for large-scale surveillance.
This work marks a significant advancement over prior research, which often highlighted the theoretical possibility of such attacks but struggled with practical deployment due to limitations in data collection and scalability. Björklund and his team have overcome these hurdles, presenting a practical, robust, and highly accurate method for monitoring entire streaming services. Their findings underscore a fundamental privacy vulnerability woven into the fabric of how we consume digital video, highlighting that even encrypted channels may not fully protect user viewing habits from determined adversaries.
The implications of this research are profound. For individuals, it means that their streaming activities, often perceived as private, could be meticulously cataloged. For nation-states or malicious actors, it opens avenues for mass surveillance, censorship evasion detection, or even targeted disinformation campaigns based on observed media consumption. The urgency for streaming services to address this inherent leak is paramount, as the current landscape leaves millions of users vulnerable to unprecedented levels of privacy invasion.
Background
▶ Watch: Introduction: Encrypted video privacy vulnerability (0:00)
Modern video streaming has evolved significantly from its early, often inefficient, UDP-based protocols like RTP. Today, the vast majority of video content delivery relies on HTTP-based Adaptive Bit Rate (ABR) streaming. This paradigm is designed for efficiency and user experience, enabling clients to dynamically adjust the video quality based on network conditions. Videos are encoded at multiple bit rates and then segmented into smaller, digestible chunks, typically ranging from 4 to 15 seconds in length. When a user initiates a stream, their client first fetches a manifest file (e.g., an .mpd for DASH or .m3u8 for HLS), which provides metadata about available qualities and the location of individual segments.
A critical component of ABR streaming, and the root of the privacy vulnerability, is the use of Variable Bit Rate (VBR) encoding. Unlike fixed bit rate encoding, VBR intelligently allocates bits based on the complexity of the video content within each segment. For instance, a static scene with a stationary speaker (low entropy) requires fewer bits to encode, resulting in a smaller file size. Conversely, a dynamic scene filled with motion, like confetti falling (high entropy), demands significantly more bits to maintain visual fidelity, leading to a larger file size.
The direct consequence of VBR encoding is that video segments are not uniform in size; instead, they are variable-sized segments. As a client downloads and plays these segments, it typically follows a bursty network pattern. A new segment is requested when the playback buffer falls below a certain threshold, leading to a sudden burst of network traffic. This burst's size precisely correlates to the size of the downloaded video segment. When observed consecutively, the sequence of these variable-sized network bursts forms a unique fingerprint for a particular video. Just as human fingerprints are unique, the sequence of segment sizes for a specific video across its duration creates a distinctive pattern that can be identified.
Prior research has acknowledged the theoretical possibility of such fingerprinting attacks. However, these works largely faced practical limitations that prevented their deployment in real-world scenarios. Common challenges included:
- Small Datasets: Relying on limited collections of videos, making large-scale identification impractical.
- Base Rate Neglect: Difficulty in distinguishing a specific video from a vast pool of possibilities without a comprehensive dataset.
- Costly Data Collection: The need to stream entire videos to collect fingerprints, which is time-consuming and resource-intensive.
- Specific Network Characteristics: Many attacks required particular network conditions or relied on features that were not universally available, limiting their applicability.
The goal of Björklund's work was to demonstrate that this attack is not merely possible, but eminently practical and scalable for monitoring entire streaming services, overcoming the shortcomings of previous efforts.
Key Findings
▶ Watch: Bursty network traffic reveals exact video content (2:15)
The research presented by Martin Björklund at USENIX Security yielded several critical findings that collectively demonstrate the severe privacy implications of modern video streaming protocols:
- Demonstrated Practicality and Scalability: The most significant finding is the definitive proof that video fingerprinting attacks are not just theoretical but are practical to deploy for monitoring entire streaming services at scale. This overcomes a major hurdle that limited the real-world applicability of previous research.
- Largest-Ever Dataset: The team amassed an unprecedented dataset for video fingerprinting. They collected approximately 240,000 videos, which translated into over 3 million unique segment fingerprints. This vast dataset is crucial for the attack's high accuracy and ability to distinguish between a multitude of videos.
- Highly Efficient Data Collection: A novel and efficient data collection method was developed. By leveraging manifest files, the researchers could obtain all necessary segment size information for an entire streaming service library in less than a day, without ever having to stream the actual video content. This efficiency ensures the attacker can maintain an up-to-date database of video fingerprints as service libraries evolve.
- Protocol Agnostic Attack: The attack's effectiveness does not hinge on specific network characteristics or protocols. Its reliance on the fundamental properties of ABR streaming makes it broadly applicable across different platforms and network environments.
- Resilience to Obfuscation: Crucially, the attack proved effective even when targets employed common privacy measures. It worked successfully against targets using a Virtual Private Network (VPN) to mask their traffic and even through passive Wi-Fi sniffing, highlighting the limitations of these tools in protecting streaming privacy.
- Exceptional Accuracy and Precision: The system achieved a remarkable 99.5% accuracy in identifying videos after just 10 minutes of eavesdropping. Furthermore, it recorded zero false positives for both standard HTTPS (strong attacker) and VPN (weak attacker) scenarios, indicating a very high level of precision and confidence in its predictions.
- Inherent Vulnerability in ABR Streaming: The research concludes that the leak of viewing information is inherent to ABR streaming itself. This suggests that simple patches or third-party workarounds are insufficient, and fundamental changes to how ABR streaming is implemented by service providers are necessary to mitigate this privacy risk.
These findings collectively paint a stark picture: the way we currently stream video creates an unavoidable side channel that can be exploited for large-scale, accurate, and persistent surveillance of user viewing habits.
Technical Deep Dive
▶ Watch: Efficient data collection via manifest files (5:45)
The core of the "Endangered Privacy" attack lies in exploiting the predictable yet unique patterns generated by Adaptive Bit Rate (ABR) streaming and Variable Bit Rate (VBR) encoding. When a video is encoded using VBR, scenes with high visual complexity (e.g., action sequences) require more data bits, resulting in larger segment files, while simpler scenes (e.g., static shots) require fewer bits, leading to smaller segment files. As a video progresses, the sequence of these variable-sized segments creates a distinctive "fingerprint." When a client streams, it downloads these segments, generating a corresponding sequence of network bursts whose sizes directly mirror the segment sizes. This sequence of burst sizes, even over an encrypted connection, becomes the observable leak.
To operationalize this, the researchers developed a sophisticated multi-stage system:
1. Data Collection Methodology:
The initial and most critical step is building a comprehensive database of video fingerprints. Unlike previous approaches that involved streaming entire videos, which is time-consuming and resource-intensive, Björklund's team devised a highly efficient method:
- They developed a dedicated program designed to parse the manifest files (e.g.,
.mpdor.m3u8) that streaming clients request at the beginning of a session. - These manifest files contain metadata detailing all available video qualities (bit rates) and, crucially, the sizes and URLs of every individual segment within the video.
- By extracting this segment size information directly from the manifest files, the program could generate full segment fingerprints for all encodings of a video without ever downloading the actual video content.
- This technique allowed them to collect fingerprints for entire libraries of major streaming services like Amazon Prime Video, Max, and the Swedish service SVT in less than a day. This efficiency is paramount for maintaining an up-to-date fingerprint database as streaming libraries frequently change. The scale of this collection was substantial, encompassing approximately 240,000 videos and generating over 3 million unique segment fingerprints.
2. Fingerprint Organization and Search:
Once collected, these raw segment fingerprints needed to be organized for rapid comparison during the live attack:
- Each full video fingerprint (a long sequence of segment sizes) was broken down into smaller, overlapping sequences. The researchers chose sequences of eight consecutive segment sizes. This transforms each sequence into an 8-dimensional data point.
- These 8-dimensional points were then stored in a KD tree (K-dimensional tree). A KD tree is a binary tree-based data structure that efficiently organizes points in a k-dimensional space, enabling very fast nearest-neighbor searches. Conceptually, it's like using a "sliding window" over a video's full fingerprint and storing each resultant 8-segment sequence in the tree.
- The sheer volume of data resulted in a KD tree containing approximately 2 billion points. Despite this massive scale, the KD tree's optimized structure allowed for quick querying, which is essential for real-time identification.
3. Live Attack Execution:
For the live monitoring phase, the attacker needs to observe the target's network traffic:
- A T-shark wrapper was developed. T-shark is the command-line version of the popular network protocol analyzer Wireshark.
- This wrapper continuously monitors the target's network interface, summing up bursts of network traffic. Each significant burst directly corresponds to the download of a single video segment by the streaming client.
- The sizes of these observed network bursts, taken in sequence, form the live "fingerprint" input for the system.
- These live 8-dimensional burst sequences are then used to query the pre-built KD tree. The tree rapidly returns a set of candidate videos whose stored fingerprints most closely match the observed live traffic pattern.
4. Identification Heuristic:
A single query to the KD tree is typically insufficient for confident identification due to potential network noise, slight timing variations, or the possibility of similar segment patterns between different videos. To address this, a sophisticated heuristic was employed:
- The core principle of the heuristic is that the correct video should appear more often and with a consistently closer distance among the candidates returned by the KD tree over time, compared to incorrect videos.
- The system continuously makes predictions, typically at each "time step" (e.g., every 4 seconds, corresponding to a new segment request).
- It calculates a confidence score for each candidate video. The video with the highest score that exceeds a predefined threshold (referred to as
theta) is then declared as the identified stream.
5. System Behavior Over Time:
The talk vividly illustrated the system's dynamic identification process. In a demonstration, a target streamed three consecutive videos from different platforms (Max, Amazon, SVT) over a 30-minute period:
- Initially, as the target streamed a video from Max, the system's confidence score for that specific Max video steadily increased. Once the score surpassed the threshold (e.g., 2.2 in the example), the system confidently identified the video.
- When the target switched to an Amazon video, the confidence score for the previous Max video began to decay, while the score for the new Amazon video rapidly rose, eventually crossing the threshold and leading to a new identification.
- This adaptive behavior was replicated when the target subsequently switched to an SVT video, demonstrating the system's ability to track and re-identify videos as a user's viewing habits change in real-time. This continuous monitoring and re-evaluation ensure high accuracy and adaptability.
Demo / Proof of Concept
▶ Watch: KD-tree for fast fingerprint matching (6:20)
To thoroughly evaluate the attack's efficacy in an authentic network environment, the researchers devised a robust experimental setup and tested it under various attacker models.
Experimental Setup:
- A dedicated Local Area Network (LAN) was established to simulate a controlled but realistic network environment.
- A man-in-the-middle (MitM) eavesdropper device was strategically positioned behind an access point, mirroring a scenario where an ISP, a malicious actor controlling public Wi-Fi, or even a sophisticated home network intruder could intercept traffic.
- Three distinct target devices were used: a Linux laptop, a Mac laptop, and a Windows laptop. These devices were all connected to the MitM-controlled access point.
- To ensure efficiency and simulate realistic user behavior at scale, an automated web driver script was deployed on each target machine. This script programmatically navigated streaming services (Amazon Prime Video, Max, SVT), selected random videos, and streamed them for 10-minute intervals before switching to a new video. This automated process allowed for the rapid generation of extensive, varied, and authentic network traffic for analysis.
Evaluated Attack Models:
- Strong Attacker (HTTPS/TLS):
- This model represents the most common eavesdropping scenario, such as an Internet Service Provider (ISP) or a sophisticated network adversary.
- The target streams video over standard HTTPS/TLS without a VPN.
- The strong attacker can easily isolate the video stream from other network traffic. This is typically achieved by inspecting the TLS handshake to read the Server Name Indication (SNI) field (e.g.,
primevideo.com,max.com), which reveals the destination streaming service. Alternatively, the attacker could maintain a lookup table of known IP addresses associated with major streaming services. - In this setup, the attack proved exceptionally effective.
- Weak Attacker (VPN):
- This model simulates a scenario where the target attempts to mask their traffic by using a Virtual Private Network (VPN) service.
- Here, the attacker cannot directly inspect SNI or easily isolate the video stream, as all traffic is encapsulated within the encrypted VPN tunnel. The attacker must contend with additional network noise generated by the VPN traffic itself and other background activities.
- The researchers made a crucial assumption: the target was primarily engaged in video watching, meaning while background traffic (ads, telemetry from the streaming app, other background applications) contributed noise, the dominant traffic pattern was still the video stream.
- Even under these more challenging conditions, the attack remained highly potent.
- Wi-Fi Setting (Passive Sniffing):
- This model involved passively sniffing network traffic over Wi-Fi, which often introduces more noise and potential packet loss compared to a wired MitM setup.
- The behavior was observed to be similar to the VPN setting in terms of dealing with noise.
- However, the researchers noted a specific limitation in their own monitoring device: it occasionally missed a significant number of frames during capturing. This issue was particularly pronounced with Amazon videos, which tend to have a very large initial buffer fill, potentially overwhelming the monitoring card. While this specific hardware limitation led to slightly worse results for Amazon videos in this particular test, it did not invalidate the underlying attack methodology.
Results:
The experimental results unequivocally demonstrated the attack's devastating effectiveness:
- For both the strong attacker (HTTPS) and the weak attacker (VPN) settings, the system achieved a remarkable 99.5% accuracy in identifying the correct video after just 10 minutes of continuous eavesdropping.
- Critically, both these scenarios recorded zero false positives, indicating that when the system identified a video, it was almost certainly correct.
- The paper also discusses the importance of addressing base rate neglect in the VPN and Wi-Fi scenarios. When dealing with the "set of all internet traffic" that a VPN or passive Wi-Fi capture might present, the potential for false positives increases if the identification heuristic isn't robust enough to filter out non-video traffic or correctly prioritize candidates. However, their heuristic proved capable of mitigating this risk effectively.
These results confirm that the segment fingerprinting attack is not only robust but also highly practical and resilient against common privacy countermeasures like VPNs, posing a significant threat to user privacy on a global scale.
Defensive Implications
▶ Watch: Real-time demonstration of video identification confidence (8:40)
The research by Martin Björklund presents a stark challenge to the current architecture of video streaming and highlights the urgent need for robust defensive measures. The core message is clear: the privacy leak is inherent to Adaptive Bit Rate (ABR) streaming itself, making traditional encryption insufficient and simple workarounds largely ineffective.
1. Traffic Manipulation (Adding Noise):
The most intuitive and direct defense against fingerprinting attacks is to manipulate the network traffic, essentially adding noise or obfuscation to break the direct correlation between segment sizes and observed network bursts.
- The researchers evaluated such a solution implemented by the VPN company Mulvad. Mulvad's approach involves actively altering the traffic patterns, and it was found to be effective in stopping the fingerprinting attack.
- However, this mitigation comes with a significant drawback: it is very costly. The Mulvad solution resulted in approximately 2.2 times more data usage for the user. This overhead is substantial, impacting users with data caps, slower internet connections, or those in regions where bandwidth is expensive.
- The speaker emphasized that expecting every user to adopt such a costly third-party solution is impractical and unsustainable. It shifts the burden of a fundamental protocol flaw onto the end-user.
2. The Inherent Leak in ABR Streaming:
The central argument is that the vulnerability isn't a bug in a specific implementation, but a fundamental byproduct of how ABR streaming and VBR encoding are designed to work efficiently. The very mechanism that allows for adaptive quality and efficient bandwidth use also creates the side channel for fingerprinting. This means:
- Standard HTTPS/TLS encryption protects the content of the video but does not hide the size or timing of the packets, which is precisely what the attack leverages.
- Unless a VPN actively modifies traffic, it merely routes the same bursty, variable-sized patterns through an encrypted tunnel, leaving the fingerprint intact.
3. Proposed Mitigation Strategy (Buffer Management):
Given that the leak is inherent to ABR, the solution must lie within the ABR implementation itself. While the transcript does not provide exhaustive technical details, the paper suggests that the mitigation strategy revolves around the way the buffer is managed by streaming clients and servers. This implies techniques that aim to decouple the observed network traffic patterns from the actual segment sizes, thereby destroying the unique fingerprint:
- Segment Padding: Streaming services could add random or fixed-size padding to smaller segments to normalize their observed network size, making all segments appear to be of a similar size.
- Randomized Fetching/Buffering: Introducing randomness in when segments are requested, or varying the buffer thresholds in a less predictable manner, could obscure the direct correlation between playback progress and network bursts.
- Pre-fetching/Over-fetching: Aggressively pre-fetching segments or downloading more segments than immediately needed, possibly even discarding some, could create a more uniform and less revealing network traffic pattern.
- Traffic Shaping at the Server: Streaming servers could actively shape the outgoing traffic to clients, smoothing out the burstiness or introducing artificial delays to obscure the segment boundaries.
Implementing such changes would require significant collaboration and investment from streaming service providers. It would need careful balancing between privacy protection and maintaining the efficiency and quality-of-experience benefits that ABR streaming offers.
4. Urgency for Streaming Services:
The talk concludes with a strong call to action, urging streaming services to address this leak as fast as possible. The scale and accuracy of the demonstrated attack mean that individual privacy is significantly endangered, and the potential for misuse (from targeted advertising to state surveillance) is too high to ignore. This highlights a responsibility for streaming platforms to re-evaluate their ABR implementations with privacy as a core design principle, rather than an afterthought.
Key Takeaways
- Video streaming services are profoundly vulnerable to large-scale monitoring, even with encrypted connections, due to the inherent segment-size fingerprinting created by Adaptive Bit Rate (ABR) streaming and Variable Bit Rate (VBR) encoding.
- An attacker can accurately identify specific videos being watched by observing the unique sequence of network burst sizes, which directly correlates to the variable sizes of video segments.
- This research demonstrates the attack's practicality and scalability, achieving 99.5% accuracy with zero false positives after just 10 minutes of eavesdropping, significantly surpassing previous theoretical limitations.
- Common privacy tools like VPNs do not offer full protection against this attack unless they actively manipulate traffic patterns, as the underlying fingerprinting mechanism remains effective through encrypted tunnels.
- Mitigation strategies involving active traffic manipulation (e.g., adding noise) are effective but incur a substantial cost, such as a 2.2x increase in data usage, making them impractical for widespread adoption.
- The privacy leak is considered inherent to the ABR streaming protocol itself, necessitating fundamental changes by streaming services, likely focusing on buffer management strategies, to truly address this vulnerability.
About the Speaker(s)
Martin Björklund is the researcher who presented the paper "Endangered Privacy: Large-Scale Monitoring of Video Streaming Services" at the USENIX Security conference. While the transcript provides limited biographical details, it indicates he is a primary contributor to this significant work, highlighting his expertise in network security and privacy research, particularly concerning video streaming technologies.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid, well-executed network privacy research that moves the video fingerprinting conversation from 'theoretically possible' to 'practically deployed at scale.' The 240k-video dataset, manifest-scraping efficiency trick, and KD-tree architecture are the real contributions here — this isn't a rehash, it's a meaningful engineering advancement over prior work. Doesn't quite hit five stars because the core attack concept (ABR fingerprinting via segment sizes) has prior art, and the defensive section is thin on actionable specifics.
Heather Calloway (CISO) — SOLID
Credible, well-executed research that closes the gap between theoretical fingerprinting attacks and practical deployment at scale. The technical contribution is real, but the talk stops short of telling the people who can actually fix this what to do — streaming service security and privacy leaders get a problem statement, not a decision framework.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)