Guardians of the Galaxy: Content Moderation in the InterPlanetary File System

Saidu Sokoto, Leonhard Balduf, Dennis Trautwein, Yiluo Wei, Gareth Tyson, Ignacio Castro, Onur Ascigil, George Pavlou, Maciej Korczyński, Björn Scheuermann, Michał Król

33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24

Overview

This talk, presented by Saidu Sokoto and a team of international collaborators, delves into the complex and often contentious realm of content moderation within the InterPlanetary File System (IPFS). As a decentralized, peer-to-peer data storage and retrieval network, IPFS presents unique challenges for managing illicit or problematic content, particularly when that content becomes easily accessible via traditional HTTP through public gateways. The research specifically investigates an existing, centralized moderation solution implemented by Protocol Labs, the primary developers of IPFS, known as the "bad bits list."

Watch on YouTube

Visual summary for Guardians of the Galaxy: Content Moderation in the InterPlanetary File System by Saidu Sokoto, Leonhard Balduf, Dennis Trautwein, Yiluo Wei, Gareth Tyson, Ignacio Castro, Onur Ascigil, George Pavlou, Maciej Korczyński, Björn Scheuermann, Michał Król
Visual summary for Guardians of the Galaxy: Content Moderation in the InterPlanetary File System by Saidu Sokoto, Leonhard Balduf, Dennis Trautwein, Yiluo Wei, Gareth Tyson, Ignacio Castro, Onur Ascigil, George Pavlou, Maciej Korczyński, Björn Scheuermann, Michał Król

Key moments

  1. 0:00 Introduction to content moderation challenges in IPFS
  2. 1:30 Understanding IPFS architecture and HTTP gateways
  3. 3:10 Key challenges for content moderation in IPFS
  4. 4:05 Introducing IPFS's 'Bad Bits List' and its function
  5. 5:00 Bad Bits List growth and content categorization (DMCA, phishing)
  6. 6:00 Analysis of Bad Bits List's reaction time to new content
  7. 7:00 Frequency of requests for blocked content, especially fishing

Guardians of the Galaxy: Content Moderation in the InterPlanetary File System

Speakers: Saidu Sokoto; Leonhard Balduf; Dennis Trautwein; Yiluo Wei; Gareth Tyson; Ignacio Castro; Onur Ascigil; George Pavlou; Maciej Korczyński; Björn Scheuermann; Michał Król

Conference: USENIX Security '24

YouTube: https://www.youtube.com/watch?v=Nl8vhqVLydw

Overview

This talk, presented by Saidu Sokoto and a team of international collaborators, delves into the complex and often contentious realm of content moderation within the InterPlanetary File System (IPFS). As a decentralized, peer-to-peer data storage and retrieval network, IPFS presents unique challenges for managing illicit or problematic content, particularly when that content becomes easily accessible via traditional HTTP through public gateways. The research specifically investigates an existing, centralized moderation solution implemented by Protocol Labs, the primary developers of IPFS, known as the "bad bits list."

The core objective of this work is to characterize the effectiveness, scope, and limitations of the bad bits list. By analyzing its evolution, content types, reaction times to new illicit material, and its adoption across various IPFS gateways, the researchers offer a comprehensive assessment of its current state. The findings highlight both the successes and significant shortcomings of a centralized moderation approach within a fundamentally decentralized ecosystem, providing critical insights for both network operators and security professionals.

The talk is significant because it addresses a fundamental tension in decentralized systems: the desire for censorship resistance versus the societal need to mitigate harmful content. As IPFS gains traction for web3 projects and distributed applications, understanding how content moderation is (or isn't) enforced is crucial for its broader adoption and for safeguarding users. The research not only quantifies the prevalence of problematic content but also exposes the disparate levels of enforcement across the IPFS ecosystem, offering a pragmatic look at the real-world implications of decentralized content hosting.

Background

▶ Watch: Introduction to content moderation challenges in IPFS (0:00)

The InterPlanetary File System (IPFS) is a distributed system designed to store and access files, websites, applications, and data. Unlike traditional web models that rely on centralized servers, IPFS employs a peer-to-peer (P2P) architecture where content is addressed by its hash, known as a Content Identifier (CID). This content-addressing mechanism ensures that retrieving content means verifying its integrity, as the CID is a cryptographic hash of the content itself.

At its core, IPFS utilizes a Kademlia Distributed Hash Table (DHT). This DHT stores mappings from CIDs to a list of "providers" – IPFS nodes that are currently hosting that specific content. When a user wants to access content, their IPFS node queries the DHT to locate providers. Once providers are identified, the content is retrieved using a protocol called BitSwap, which facilitates peer-to-peer exchange of data blocks.

While IPFS is inherently decentralized, a significant portion of its accessibility to the broader public comes through HTTP Gateways. These gateways, such as ipfs.io operated by Protocol Labs, act as bridges, allowing users to access IPFS content directly through a standard web browser or any HTTP client without needing to run a dedicated IPFS node. A typical gateway URL might look like https://ipfs.io/ipfs/<CID>. The convenience of these gateways has led to the rise of Gateway-as-a-Service providers, which allow projects (e.g., web3 games) to rent dedicated subdomains to serve IPFS-hosted content via custom, often benign-looking URLs. This ease of access, however, introduces a critical challenge: it makes the dissemination of problematic content, such as phishing pages or malware, remarkably simple, as users only need to click a link.

The decentralized nature of IPFS inherently complicates content moderation. There is no central registry for CIDs, meaning anyone can add any content they wish to the network. The public gateways, while convenient, make it difficult to determine the ultimate origin of content, as requests are proxied through the gateway. Furthermore, due to content addressing, multiple URLs can point to the same underlying data, making traditional URL-based filtering mechanisms like Google Safe Browsing largely ineffective. These challenges underscore the need for dedicated content moderation strategies within the IPFS ecosystem.

Key Findings

▶ Watch: Key challenges for content moderation in IPFS (3:10)

The research meticulously characterizes Protocol Labs' "bad bits list," a centralized effort initiated in 2021 to address the growing problem of problematic content on IPFS. By the time of the study, the list had grown to approximately 400,000 entries.

List Mechanism and Evolution: The bad bits list is populated based on email reports sent to Protocol Labs. Upon receiving a takedown notice, Protocol Labs reviews the legitimacy of the request. If validated, the reported CIDs are normalized (to account for different encodings), hashed, and then added to the public list. Recent versions of IPFS software include support for checking incoming requests against this list, denying access if a CID is found on it. The list exhibited significant growth, particularly in 2023, with large batches of entries being added, indicating a reactive response to an escalating problem.

Content Characterization: A detailed analysis of the list's contents revealed that the vast majority of entries are related to DMCA (copyrighted material). These often manifest as PDFs, suggesting a prevalence of pirated research papers or similar documents. Beyond DMCA, the list also includes entries for phishing content (primarily HTML files), malware, and other unspecified problematic material. A portion of content remained uncharacterized due to an inability to download or access it, indicating the ephemeral nature of some distributed content.

Reaction Time and Responsiveness: The study examined the latency between content appearing on the IPFS network and its subsequent inclusion in the bad bits list. The overall reaction time was found to be "quite slow," primarily because the list is passive and relies on active email reports. However, once Protocol Labs receives an email report, their internal reaction time to process and add the item to the list is "quite quickly." This contrasts with anti-fishing services (like OpenPhish or Google Safe Browsing), which generally react much faster, suggesting a more proactive or automated detection mechanism.

Requests for Blocked Content: Despite the moderation efforts, the researchers observed that less than 1% of total requests seen across the network were for content currently on the bad bits list. While this percentage might seem small, the volume of requests for problematic content is significant. Copyrighted material consistently remained popular, but a notable and concerning trend was the explosive increase in requests for phishing content starting around 2023, a trend that appeared to be continuing.

Gateway Implementation and Enforcement: A critical finding was the highly varied and often inconsistent implementation of the bad bits list by public IPFS gateway providers. Gateway-as-a-Service providers were found to "not care a whole lot about any of the lists," showing minimal adherence to both the bad bits list and general web2 anti-fishing services. In contrast, gateways operated directly by Protocol Labs unsurprisingly adhered strictly to their own bad bits list. The overall inconsistency stems from the fact that implementing the list is optional, requires running the latest software versions, subscribing to list updates, and actively maintaining it. Consequently, many clients and gateways simply do not implement the list, meaning that content flagged as problematic remains widely available through the peer-to-peer network and through non-compliant public gateways.

Technical Deep Dive

▶ Watch: Introducing IPFS's 'Bad Bits List' and its function (4:05)

The technical underpinnings of IPFS and the bad bits list reveal a fascinating interplay between decentralized design principles and centralized security interventions. At its core, IPFS functions by content addressing, where each piece of data is identified by a unique Content Identifier (CID), which is a cryptographic hash of the data itself. This means that if even a single bit of the content changes, its CID changes, ensuring data integrity and immutability.

When content is added to an IPFS node, that node becomes a "provider" for that content. It then announces its provision to the network's Kademlia Distributed Hash Table (DHT). The Kademlia DHT is a key-value store where CIDs act as keys, and the values are lists of peer IDs (network addresses) of nodes providing that content. Retrieving content involves a multi-step process: a client queries the DHT for a given CID, obtains a list of providers, and then connects directly to one or more providers to download the content blocks using the BitSwap protocol. BitSwap is a peer-to-peer data exchange protocol designed for efficiency in a distributed environment.

HTTP gateways provide a crucial abstraction layer. When a request for an IPFS CID comes to an HTTP gateway (e.g., ipfs.io/ipfs/<CID>), the gateway acts as an IPFS client on behalf of the user. It resolves the CID via the DHT, retrieves the content from providers using BitSwap, and then serves it back to the user over HTTP. This seamless integration makes IPFS content accessible to anyone with a web browser, bypassing the need for specialized IPFS software.

The "bad bits list" introduces a centralized security mechanism into this decentralized ecosystem. Maintained by Protocol Labs, it is a list of CIDs that have been identified as problematic (e.g., DMCA violations, phishing, malware). The list is publicly available, allowing any IPFS client or gateway operator to download and integrate it. The moderation process is reactive:

  1. Reporting: Individuals or organizations report problematic content (identified by its CID) via email to Protocol Labs.
  2. Review and Normalization: Protocol Labs staff review the report for legitimacy. If valid, the reported CIDs are normalized to a consistent format (as CIDs can have different encodings).
  3. Hashing and Inclusion: The normalized CIDs are then hashed (likely for efficient lookup) and added to the bad bits list.

For IPFS software (both standalone clients and gateway implementations) that supports the bad bits list, the process for an incoming request is as follows:

  1. CID Extraction: The CID from the incoming request is extracted.
  2. Normalization and Hashing: The CID is normalized and hashed, mirroring the process used when adding entries to the bad bits list.
  3. List Lookup: The hashed CID is checked against the local copy of the bad bits list.
  4. Access Control: If the CID is found on the list, the request is denied, and the content is not served. If not found, the request proceeds, and the content is retrieved from the IPFS network.

A significant technical challenge highlighted by the research is the lack of annotations within the bad bits list. The list simply contains CIDs without metadata indicating the type of problematic content (e.g., DMCA, phishing, malware). This absence of context forces gateway operators to either implement the entire list or none of it. For instance, an operator in a jurisdiction where DMCA violations are not a priority might choose to ignore the entire list, inadvertently allowing access to malware or phishing content that they would want to block. This "all or nothing" approach hinders nuanced and localized moderation strategies.

Furthermore, the optional nature of list implementation, coupled with the need for operators to run up-to-date software and actively subscribe to list updates, leads to a fragmented enforcement landscape. While Protocol Labs' own gateways demonstrably adhere to the list, many third-party and Gateway-as-a-Service providers do not, effectively creating "safe havens" for illicit content accessible via HTTP. This exposes a fundamental tension: a centralized moderation list in a decentralized system is only as effective as its decentralized adoption and enforcement.

Demo / Proof of Concept

▶ Watch: Analysis of Bad Bits List's reaction time to new content (6:00)

The talk did not feature a traditional "demo" in the sense of a live demonstration of a new tool or exploit. Instead, the research itself served as a comprehensive characterization and analysis of an existing content moderation system within IPFS. The "demonstration" of their findings was presented through compelling data visualizations and statistical analysis derived from their extensive measurements of the IPFS network and the bad bits list.

Specifically, the speakers used figures and graphs to illustrate:

  • The explosive growth of the bad bits list: Showing how the number of entries, particularly in 2023, reflected increasing moderation efforts.
  • The breakdown of content types: Visually representing the dominance of DMCA content (e.g., PDFs) versus phishing (e.g., HTML) and other categories.
  • Reaction times: Comparing the overall time from content appearance to list inclusion against Protocol Labs' internal response time and general anti-fishing services, highlighting the bottlenecks.
  • Request patterns for blocked content: Showing how copyrighted material remains consistently popular, while phishing content saw a significant surge in requests.
  • Gateway implementation disparities: A complex figure illustrated which types of gateway operators (e.g., Gateway-as-a-Service, Protocol Labs) adhered to which blocklists (bad bits list, web2 anti-fishing lists), visually confirming the inconsistent enforcement.

These data-driven insights served as the "proof of concept" for their claims regarding the state of content moderation in IPFS. They effectively demonstrated the operational realities of the bad bits list and the varied compliance across the network, providing empirical evidence for the challenges and limitations discussed.

Defensive Implications

▶ Watch: Frequency of requests for blocked content, especially fishing (7:00)

The findings presented in "Guardians of the Galaxy" offer several critical defensive implications for various stakeholders within and around the IPFS ecosystem.

For IPFS Gateway Operators, the message is clear: implementing and consistently updating the bad bits list is a crucial first step in responsible content moderation. The research highlights the significant disparity in enforcement, particularly among Gateway-as-a-Service providers who often disregard moderation lists entirely. Operators should:

  1. Prioritize Implementation: Ensure their gateway software is up-to-date and configured to check against the latest version of the bad bits list.
  2. Subscribe to Updates: Establish mechanisms to regularly fetch and apply updates to the bad bits list to ensure ongoing protection against newly identified problematic content.
  3. Consider Broader Filtering: While the bad bits list is IPFS-specific, integrating with established Web2 anti-fishing and malware lists (as some gateways do) can provide an additional layer of defense.

For Users of IPFS and Web3 Applications, it's essential to understand the limitations of current moderation efforts. Even if content is blocked by some public gateways, it remains available on the underlying peer-to-peer IPFS network. Users running their own IPFS nodes can still access this content directly. Therefore:

  1. Exercise Caution: Do not assume that content accessed via IPFS links is inherently vetted or safe, especially if the link comes from an untrusted source or uses a Gateway-as-a-Service provider.
  2. Understand Node Behavior: Be aware that running a full IPFS node bypasses gateway-level moderation, granting access to all content on the network.

For Protocol Labs and the broader IPFS Community, the research points towards several areas for improvement to foster a more robust and decentralized moderation framework:

  1. Decentralized Moderation Solutions: The current centralized bad bits list, while effective where implemented, runs counter to the decentralized ethos of IPFS. Research and development into genuinely decentralized, community-driven moderation solutions are paramount. This could involve reputation systems, cryptographic attestations, or distributed consensus mechanisms for content flagging.
  2. Annotated Lists: The lack of annotations in the bad bits list (e.g., distinguishing DMCA from malware) forces an "all or nothing" adoption. Adding metadata to list entries would allow gateway operators to implement more granular, jurisdiction-aware, and threat-specific blocking policies. This would encourage broader adoption by allowing operators to selectively block content types relevant to their operational context.
  3. Improved Interoperability: Enhancing the ability for Web3 filtering lists to ingest and leverage data from established Web2 anti-fishing and malware services would significantly improve the speed and coverage of moderation efforts. Conversely, Web2 filtering services could be improved to understand and filter CIDs more effectively.
  4. Proactive Detection: The slow reaction time for new content to appear on the bad bits list (due to its reliance on email reports) suggests a need for more proactive detection mechanisms, perhaps leveraging AI/ML or community-driven reporting tools that can rapidly identify and flag problematic CIDs.

Finally, Threat Actors should recognize that while IPFS offers a resilient platform for hosting content, the visibility and moderation efforts are increasing. While inconsistencies in gateway enforcement may provide temporary havens, the centralized bad bits list is actively growing and targets problematic CIDs. The surge in phishing requests, in particular, indicates that IPFS is being exploited, but this exploitation is also drawing attention and resources towards developing more effective countermeasures.

Key Takeaways

  • IPFS content moderation is complex due to its decentralized nature and the accessibility provided by HTTP gateways. The tension between censorship resistance and the need to block harmful content is a core challenge.
  • Protocol Labs' "bad bits list" is a centralized moderation effort containing over 400,000 entries. The majority of listed content relates to DMCA violations, but phishing and malware are also significant concerns, with phishing requests seeing a sharp increase in 2023.
  • The reaction time for new problematic content to appear on the bad bits list is generally slow, as it relies on email reports, though Protocol Labs responds quickly once a report is received. This contrasts with faster-reacting Web2 anti-fishing services.
  • Implementation of the bad bits list across IPFS gateways is highly inconsistent and optional. Gateway-as-a-Service providers often show minimal adherence, while Protocol Labs' own gateways fully implement it, leading to a fragmented enforcement landscape.
  • Blocked content remains accessible directly via the peer-to-peer IPFS network. Gateway moderation only restricts access through specific HTTP entry points, not the underlying distributed storage.
  • Future improvements require decentralized moderation solutions, better list annotations, and enhanced interoperability between Web2 and Web3 filtering mechanisms to achieve more comprehensive and effective content moderation.

About the Speaker(s)

The primary presenter of this talk was Saidu Sokoto, who delivered the research findings on behalf of a large, international collaborative team. The extensive list of co-authors includes Leonhard Balduf, Dennis Trautwein, Yiluo Wei, Gareth Tyson, Ignacio Castro, Onur Ascigil, George Pavlou, Maciej Korczyński, Björn Scheuermann, and Michał Król. Saidu Sokoto expressed gratitude to his co-authors, emphasizing that the work would not have been possible without their collective effort. The collaboration across multiple institutions highlights the global interest and complexity of addressing content moderation challenges in decentralized systems like IPFS.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This research meticulously characterizes the effectiveness and failures of Protocol Labs' "bad bits list" in moderating content on IPFS. It highlights a critical tension between decentralization and the practical necessity of content control, exposing significant inconsistencies in gateway enforcement and a concerning surge in phishing content. The findings provide vital, data-driven insights for securing decentralized systems.

Heather Calloway (CISO) — STRONG ACCEPT

This research provides a clear, unsentimental assessment of content moderation failures within IPFS. It starkly highlights the real-world business exposure from inconsistent gateway enforcement, particularly the surge in phishing content, demanding executive attention for any organization engaging with decentralized systems.

→ Top-rated talks at 33rd USENIX Security Symposium

All talks from 33rd USENIX Security Symposium