Duumviri: Detecting Trackers and Mixed Trackers with a Breakage Detector
He Shuang
Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Web Security
Overview
In the realm of digital privacy, the pervasive issue of online tracking continues to challenge users and developers alike. While content blockers have become a staple for many, their reliance on manually curated filter lists presents significant limitations, particularly in scalability and the potential for human error. This talk, presented by He Shuang at the NDSS Symposium, introduces Duumviri, a novel approach to automated tracker detection that fundamentally rethinks the problem. Duumviri distinguishes itself by proposing a two-model system: a traditional tracker detector paired with an innovative breakage detector.
Key moments
- 0:00 Introduction and problem: high webpage breakage
- 2:00 Two reasons for breakage, including mixed trackers
- 3:28 Duumviri's two-model approach with breakage detector
- 4:40 Defining breakage and collecting samples for training
- 6:00 Key methodology: reconstructing breakages from exception rules
- 7:00 How Duumviri reconstructs a breakage example
- 8:05 Large-scale evaluation and accuracy results
Duumviri: Detecting Trackers and Mixed Trackers with a Breakage Detector
Speakers: He Shuang
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=pJnVztzU7kQ
Overview
In the realm of digital privacy, the pervasive issue of online tracking continues to challenge users and developers alike. While content blockers have become a staple for many, their reliance on manually curated filter lists presents significant limitations, particularly in scalability and the potential for human error. This talk, presented by He Shuang at the NDSS Symposium, introduces Duumviri, a novel approach to automated tracker detection that fundamentally rethinks the problem. Duumviri distinguishes itself by proposing a two-model system: a traditional tracker detector paired with an innovative breakage detector.
The core motivation behind Duumviri is to overcome the critical flaw of previous automated tracker detection systems: their unacceptably high rate of breaking legitimate website functionality. Prior research, despite its advancements in identifying trackers through sophisticated graph-based analysis of web page rendering, often renders a substantial percentage of pages unusable—an issue the speaker highlights as making such tools impractical for real-world deployment. Duumviri directly confronts this challenge by explicitly incorporating a mechanism to identify and prevent such breakages, thereby enabling a more aggressive and accurate identification of trackers without compromising the user experience.
This work is particularly significant because it addresses a fundamental trade-off in privacy protection: the balance between blocking unwanted tracking and maintaining website functionality. By introducing a dedicated breakage detector, Duumviri not only promises to enhance the efficacy of tracker detection but also paves the way for a new generation of privacy tools that are both powerful and user-friendly. Furthermore, the system's ability to tackle the complex problem of mixed trackers—requests that blend functional and tracking information—marks a crucial step forward in adapting to the evolving tactics of online tracking.
Background
▶ Watch: Introduction and problem: high webpage breakage (0:00)
The landscape of online privacy protection has long been dominated by content blockers that operate on the principle of filter lists. Tools like EasyList and EasyPrivacy, widely adopted by users, scrutinize network requests against extensive, human-written rules. If a request matches a rule, it is blocked, thereby preventing trackers from loading or executing. While effective to a degree, this traditional approach faces inherent challenges. Firstly, scalability is a major concern; as new tracking methods emerge, privacy developers must continually investigate, analyze, and manually update these rules. This reactive process is labor-intensive and often struggles to keep pace with the dynamic nature of online tracking. Secondly, being human-written, these filter lists are susceptible to human error, occasionally leading to the blocking of legitimate, non-tracking requests. Such misidentifications can inadvertently disrupt website functionality, a phenomenon known as "breakage."
Recognizing these limitations, the research community has dedicated significant effort to developing automatic tracker detectors. Early attempts drew features from various sources, but more recent and sophisticated approaches have treated web page rendering as a complex graph. These models analyze interactions between entities such as HTML elements, network requests, and JavaScript executions, extracting features from this intricate graph to predict whether a given request constitutes a tracker. While these methods represent a considerable leap forward in automation, they have consistently fallen short in one critical aspect: usability.
He Shuang emphatically states that previous automated tracker detectors suffer from a high percentage of web page breakages, with some tools breaking as many as 15% of pages. This level of disruption renders such tools "really not usable" in practice. Duumviri identifies two primary reasons behind this pervasive breakage issue:
- Misidentification of Functional Requests: The root cause here lies in the training data. Previous automated detectors are often trained using existing, human-written filter lists as their ground truth. As previously noted, these filter lists are imperfect; they can contain errors where functional requests are erroneously classified as trackers. Consequently, models trained on this imperfect data inherit and amplify these errors. The speaker highlights the inherent difficulty in improving the quality of these foundational datasets, as "skilled workers" are already dedicated to refining them daily. When a functional request is misidentified and subsequently blocked, the integrity of the web page is compromised, leading to a breakage.
- Mixed Trackers: A more insidious and complex problem arises when functional information is interwoven with tracking-related information within a single request. These are termed mixed trackers. The talk provides a clear example: a request sent after a user clicks an item on a search page. This request contains multiple parameters: the first two are essential for redirecting the user to the correct item page, ensuring functionality. However, a third parameter contains a high-entropy identifier, which is clearly for tracking purposes. If this entire request is blocked, the user's action fails, and the page becomes unresponsive—a breakage. Conversely, allowing the request compromises user privacy due to the transmission of the tracking identifier. This dilemma underscores a critical challenge that current content blockers and previous automated detectors struggle to resolve, as they often lack the granularity to differentiate between functional and tracking components within a single request. Duumviri aims to explicitly address this problem by providing a mechanism to detect and potentially mitigate mixed trackers.
Key Findings
▶ Watch: Duumviri's two-model approach with breakage detector (3:28)
Duumviri's most significant contribution is its innovative two-modeled approach to tracker detection, designed to overcome the high breakage rates plaguing prior automated systems. At its core, Duumviri operates with two distinct machine learning models:
- Tracker Detector: This model functions like traditional tracker detectors, answering the fundamental question: "Is a given request a tracker?"
- Breakage Detector: This is Duumviri's novel component, addressing a critical usability concern: "Does blocking a given request break the page?"
A request is ultimately classified as a tracker and subsequently blocked only if both conditions are met: the tracker detector identifies it as a tracker, and the breakage detector confirms that blocking it will not break the page. This dual-validation mechanism represents a paradigm shift, as it integrates usability directly into the detection process. The speaker notes that cases where both detectors are positive (a tracker that causes breakage if blocked) often indicate mixed trackers, a topic explored in greater detail in the accompanying paper.
The primary benefit of integrating the breakage detector is its ability to allow the tracker detector to be tuned much more aggressively. Without the fear of causing breakages, the tracker detector can be optimized to identify more potential trackers. When a misidentification of a functional request occurs, the breakage detector acts as a safety net, "pulling back" the erroneous block and thereby increasing the overall tracker detection accuracy and, critically, the usability of the system.
A crucial aspect of developing the breakage detector was defining what constitutes a "breakage." Duumviri defines a breakage as significant changes applied to the "vanilla unbroken version of the page," changes that are substantial enough for a human user to perceive as a functional disruption. To detect these, the breakage detector utilizes differential features, comparing the state of a potentially broken page to its original, functional counterpart.
Collecting training data for the breakage detector presented a significant hurdle. Breakages are inherently rare on the live web, and artificially crafting them could introduce bias. Duumviri ingeniously addresses this by studying exception rules found in existing filter lists. These rules are particularly valuable because privacy developers insert them specifically to fix reported breakages. For example, an exception rule like @@||abc.com^$domain=example.com indicates that requests to abc.com should not be blocked when on example.com, because blocking it previously caused a breakage.
Duumviri reconstructs breakages from these exception rules by visiting the specified site (e.g., example.com) and, instead of applying the exception rule (allowing the resource), it applies a "flipped rule". This flipped rule actively blocks the exact resource (e.g., abc.com) that the exception rule was designed to allow. By blocking this critical resource, Duumviri can reliably reconstruct the breakage that privacy developers previously identified and fixed. This method has several advantages: it leverages a vast existing dataset of "known" breakages, and these breakages are highly relevant to real-world scenarios. The efficacy of this reconstruction method was validated, with Duumviri successfully matching 100% of user-reported breakage cases, significantly outperforming previous work that only managed around 85%.
The effectiveness of Duumviri was demonstrated through a large-scale tracker evaluation conducted on 15,000 web pages. Using existing filter lists as a ground truth, Duumviri achieved an impressive 96.53% accuracy. Recognizing that filter lists themselves are imperfect, a disagreement analysis was performed. In cases where Duumviri's tracker detector identified a request as a tracker while the filter list did not, Duumviri was found to be correct in 55% of those cases, indicating its ability to identify trackers missed by human-curated lists.
Furthermore, Duumviri proved capable of detecting filter list-caused breakages. An example cited in the talk illustrates this: with EasyPrivacy enabled, the "entire page body is missing from the user interface," a breakage that Duumviri could identify and which had been previously reported and addressed by privacy developers. When compared to previous automated tracker detection work, Duumviri showed "extremely similar accuracy" when using filter lists as ground truth. However, in instances where the two tools disagreed on a classification, Duumviri was correct in 70.5% of the time, highlighting its superior discernment.
Finally, the talk briefly mentions Duumviri's capability to automatically detect mixed trackers, a complex challenge where functional and tracking data are combined within a single request. While not explored in depth during the presentation due to time constraints, this functionality is detailed in the full paper. The project also successfully underwent an artifact evaluation, earning the "reproduced" badge, and all code, models, and data are publicly available.
Technical Deep Dive
▶ Watch: Defining breakage and collecting samples for training (4:40)
Duumviri's architecture is centered around its two interdependent machine learning models: the Tracker Detector and the Breakage Detector. This section delves into the technical specifics of how these models are designed, trained, and integrated to achieve robust and usable tracker detection.
The Tracker Detector component, while not the primary focus of the talk, is responsible for identifying requests that carry tracking intent. The speaker, during the Q&A, reveals that this detector leverages flow-based features. This implies a sophisticated analysis of how data moves within the web page environment. Specifically, if there is a detected data flow from known "data sources" (such as cookie jars, which store user identifiers and session information) to "sinks" (primarily the network, indicating data being transmitted off the user's device), the detector extracts features from this flow. This approach is considered highly robust because, fundamentally, "if you have to conduct tracking, you have to send information somehow." This method is designed to be resilient against evasion attempts, as trackers ultimately must transmit data to fulfill their purpose. The talk doesn't detail other features used by the tracker detector, but previous work in this domain often incorporates features derived from HTML elements, JavaScript execution contexts, and the overall network request graph, suggesting Duumviri likely builds upon a comprehensive set of such indicators.
The true innovation lies in the Breakage Detector. This is a dedicated machine learning model trained to identify whether blocking a specific resource will result in a significant degradation of a web page's functionality or appearance.
Its primary function is to act as a crucial gatekeeper, ensuring that only non-breaking tracker blocks are enforced.
The design of the Breakage Detector hinges on two key elements: the definition of a breakage and the method for collecting training data.
- Defining Breakage with Differential Features:
A breakage is precisely defined as "changes that you apply to the vanilla unbroken version of the page," which must be "significant enough for a human user to call it a breakage." To quantify this, Duumviri models breakage using differential features. This involves comparing two states of a web page:
- The vanilla unbroken page: The page as it loads normally, without any blocking applied, or with necessary exceptions enabled to ensure full functionality.
- The changed page: The page after a specific request (or set of requests) has been blocked.
The differential features capture the discrepancies between these two states. While the talk doesn't list specific feature types, common differential features in web page analysis for breakage detection could include:
- DOM (Document Object Model) changes: Differences in the number of elements, their structure, attributes, or content. For example, a missing
<div>or an empty<iframe>. - Visual changes: Discrepancies in rendered pixels, layout shifts, or missing images/elements. This might involve screenshot comparisons or analysis of CSS properties.
- Network activity changes: Unfulfilled requests, long loading times for certain resources, or complete absence of expected network traffic.
- User interaction changes: Whether buttons are clickable, forms are submittable, or dynamic content loads correctly. This implies a level of functional assessment.
By training on these differential features, the model learns to distinguish between minor, insignificant changes and those that genuinely impair user experience.
- Collecting Breakage Samples from Exception Rules:
The challenge of acquiring sufficient and realistic training data for breakages is solved through an ingenious approach centered on exception rules found in existing filter lists (e.g., EasyList, EasyPrivacy). These rules are invaluable because they represent known, human-identified instances where a block caused a problem, and an exception was needed to restore functionality.
The process, as described, works as follows:
- Identification of an Exception Rule: Duumviri starts by parsing filter lists to identify exception rules. An example given is
@@||abc.com^$domain=example.com. This rule signifies that whileabc.commight generally be blocked as a tracker, it should not be blocked when the user is onexample.combecause blocking it onexample.comcauses a breakage. - Site Visitation: Duumviri then visits the website specified in the exception rule's domain (e.g.,
example.com). - Flipped Rule Application: Crucially, Duumviri does not apply the exception rule. Instead, it applies a "flipped rule". This flipped rule explicitly blocks the exact resource (e.g.,
abc.com) that the original exception rule was designed to allow. - Breakage Reconstruction: By blocking this critical resource, Duumviri effectively "reconstructs" the breakage that the original exception rule was intended to fix. This "broken" page state is then compared to a "vanilla" (unbroken) version of the page (e.g., the page loaded with the exception rule active, or the page loaded without any blocking that would cause this specific breakage).
- Feature Extraction: The differential features are then extracted from this comparison, providing a labeled sample for the breakage detector: "this set of differential features corresponds to a breakage."
This method ensures that the training data for the breakage detector is both abundant and highly relevant to real-world scenarios, directly addressing the problem of misidentified functional requests that lead to usability issues. The success of this approach is evident in Duumviri's ability to match 100% of user-reported breakages during evaluation.
Demo / Proof of Concept
▶ Watch: How Duumviri reconstructs a breakage example (7:00)
While the talk did not feature a live, interactive demonstration of Duumviri's capabilities, the speaker provided compelling evidence and evaluation results that serve as a robust proof of concept for the system's effectiveness and its underlying methodologies. The presentation included:
- Large-Scale Tracker Evaluation: Duumviri was put through a rigorous evaluation on 15,000 web pages. This extensive testing environment allowed for a statistical assessment of its performance against established benchmarks. The reported 96.53% accuracy (when using existing filter lists as ground truth) indicates a high degree of precision in identifying trackers.
- Disagreement Analysis: A critical part of the proof of concept involved a disagreement analysis. This is crucial because filter lists, used as ground truth, are themselves imperfect. When Duumviri's tracker detector identified a request as a tracker but the filter list did not, Duumviri was found to be correct in 55% of those cases. This highlights Duumviri's ability to uncover trackers that human-curated lists miss, demonstrating its superior detection capabilities in challenging scenarios.
- Demonstration of Filter List-Caused Breakages: The talk visually illustrated the very problem Duumviri aims to solve by showing an example of a filter list-caused breakage. The speaker described a scenario where, with EasyPrivacy (a popular filter list) enabled, "the entire page body is missing from the user interface." This clear and impactful example underscores how existing privacy tools can inadvertently break functionality, directly validating the need for Duumviri's breakage detector. The fact that this specific breakage had been reported and addressed by privacy developers further validates the realism of Duumviri's approach to learning from exception rules.
- Comparison with Previous Work: The proof of concept also included a direct comparison with prior automated tracker detection tools. While Duumviri achieved "extremely similar accuracy" when measured against filter lists, its strength became apparent in situations where the tools disagreed. In these cases, Duumviri was correct in 70.5% of the time, providing strong evidence of its improved discernment and reduced false positives (or false negatives, depending on the disagreement type) compared to its predecessors.
- Artifact Evaluation Badge: The successful completion of an artifact evaluation, resulting in the "reproduced" badge, serves as an independent validation of Duumviri's scientific rigor. This signifies that the research claims are verifiable, and the code, models, and data are sufficiently documented and functional for others to reproduce the stated results. The availability of all these resources at a public URL further solidifies the transparency and verifiability of the work.
While a live, interactive demo was not the format for this presentation, the comprehensive evaluation, detailed analysis of disagreements, illustrative examples of real-world problems, and independent artifact validation collectively provide a compelling proof of concept for Duumviri's innovative approach and its practical utility in enhancing both privacy and usability.
Defensive Implications
▶ Watch: Large-scale evaluation and accuracy results (8:05)
Duumviri's innovative two-model approach to tracker detection carries significant defensive implications for various stakeholders in the digital ecosystem, from individual users to privacy tool developers and even website operators.
For Users:
The most direct implication for users is the promise of more effective and less disruptive content blockers. By explicitly addressing the problem of web page breakages, Duumviri paves the way for privacy tools that can be tuned more aggressively to block trackers without sacrificing usability. This means users can benefit from enhanced privacy protection without the frustration of broken websites, potentially leading to wider adoption and sustained use of privacy-enhancing technologies. The ability to detect and mitigate mixed trackers also means users can be protected from subtle tracking mechanisms that are currently difficult for traditional blockers to handle without breaking functionality.
For Privacy Tool Developers and Researchers:
Duumviri offers a novel methodology for building next-generation tracker detectors. The concept of a breakage detector can be adopted and integrated into existing or new privacy tools, allowing developers to:
- Tune tracker detectors more aggressively: Developers can now optimize their tracker detection models for maximum coverage, knowing that the breakage detector will act as a safeguard against functional disruptions.
- Generate high-quality training data for usability: The ingenious method of reconstructing breakages from exception rules in filter lists provides a scalable and relevant way to create training datasets for usability concerns. This is a significant advancement, as collecting such data has historically been challenging.
- Automatically identify and correct filter list errors: Duumviri's capacity to detect filter list-caused breakages can empower privacy developers to automatically audit and refine their existing filter lists, improving their accuracy and reducing unintended side effects.
- Extend the concept to other security domains: The speaker explicitly mentions that the breakage detector concept could be beneficial for "other tasks where usability is a concern such as phishing detection." In phishing, blocking legitimate user interaction (e.g., login forms) while trying to block malicious elements is a common challenge. A breakage detector could help differentiate between a broken legitimate site and a malformed phishing page, or ensure that legitimate elements on a potentially malicious page are not accidentally blocked.
For Website Developers and Online Services:
Duumviri highlights the importance of separating concerns in web development. The detection of mixed trackers serves as a strong signal for website developers to:
- Decouple functional and tracking requests: To avoid being flagged and potentially blocked by sophisticated detectors like Duumviri, website operators should ensure that requests essential for core functionality do not carry tracking-related parameters or identifiers. This promotes better privacy practices and ensures their services remain accessible to users employing advanced privacy tools.
- Adapt to evolving privacy demands: The development of tools like Duumviri signals a growing trend towards more intelligent and user-centric privacy protection. Websites that fail to adapt their tracking practices may find their services increasingly disrupted for privacy-conscious users.
In the "Arms Race" between Advertisers and Ad Blockers:
The Q&A session directly addressed how Duumviri might shape this ongoing "arms race." The speaker believes that Duumviri's tracker detector, relying on flow-based features (tracking data flow from cookie jars to the network), is quite robust because tracking fundamentally requires data transmission. While advertisers might attempt to evade the breakage detector by making minor, non-functional changes appear as significant breakages, the speaker is confident that Duumviri's breakage detector has been "tuned enough" to discern genuine breakages from superficial ones. Ultimately, the speaker suggests that if a server is "willing to sacrifice functionality to enable tracking," it might resort to forcing users to choose between disabling their ad blocker or foregoing the service. However, this is a risky strategy for service providers, as users typically prioritize functionality alongside privacy. Duumviri shifts the power dynamic by making it harder for trackers to hide within legitimate functionality.
In summary, Duumviri provides a robust framework that empowers defenders with more accurate and usable tools, encourages better privacy design from website developers, and contributes a reusable methodology for addressing usability concerns across various security applications.
Key Takeaways
- High breakage rates make current automated tracker detectors unusable: Previous attempts at automated tracker detection, despite their sophistication, frequently break legitimate website functionality (up to 15% of pages), rendering them impractical for real-world application.
- Duumviri introduces a novel two-model approach: It combines a traditional tracker detector with a unique breakage detector to ensure that blocking trackers does not disrupt website functionality.
- The breakage detector is critical for usability and accuracy: By explicitly identifying when blocking a request would cause a breakage, Duumviri allows its tracker detector to be tuned more aggressively, leading to higher overall tracker detection accuracy while maintaining a seamless user experience.
- Breakage detection is trained using "flipped" exception rules: Duumviri ingeniously reconstructs real-world breakages by applying "flipped rules" derived from existing filter list exception rules, providing a scalable and relevant source of training data for its breakage detector.
- Duumviri significantly outperforms prior work in practical scenarios: It successfully matches 100% of user-reported breakages, achieves high accuracy on 15k pages (96.53%), and proves superior in disagreement analysis against both filter lists (55% correct) and previous tools (70.5% correct).
- The breakage detector concept has broader applicability: The methodology developed for detecting web page breakages can be extended to other security domains where usability is a concern, such as phishing detection, offering a versatile tool for enhancing user trust and security.
- Duumviri can detect and highlight mixed trackers: The system can identify requests that blend both functional and tracking information, a complex challenge for traditional blockers, prompting better design practices for website developers.
About the Speaker(s)
The talk was presented by He Shuang. Based on the transcript, He Shuang is the primary researcher and presenter for Duumviri. While the specific institutional affiliation was not mentioned during the presentation, the detailed technical content and the rigor of the evaluation suggest a background in computer science and security research. The presentation demonstrated a deep understanding of web tracking mechanisms, machine learning applications in security, and the practical challenges of deploying privacy-enhancing technologies.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Duumviri is legitimate academic security research solving a real problem — high breakage rates in automated tracker detection — with a clever two-model architecture and a genuinely novel training data strategy. Solid NDSS-tier systems work, but it's niche enough and incremental enough that it fills a conference slot without threatening to redefine anyone's practice.
Heather Calloway (CISO) — WEAK
Duumviri is methodologically sound academic work that solves a real engineering problem — high breakage rates in automated tracker detection — with a clever two-model design. But it never leaves the research lab. There is no path from this paper to a security program, a policy decision, or a board conversation.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025