LLFuzz: An Over-the-Air Dynamic Testing Framework for Cellular Baseband Lower Layers

Tuan Dinh Hoang (PhD student · CEC lab, Korea)

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Network Security 3: BLE and Cellular

Overview

This talk introduces LLFuzz, an innovative over-the-air dynamic testing framework designed to uncover memory corruption vulnerabilities within the lower layers of cellular basebands. Presented by Tuan Dinh Hoang, a PhD student from CEC Lab, KAIST Korea, LLFuzz addresses a critical gap in existing security research, which has predominantly focused on the higher layers (Layer 3) of cellular protocol stacks. The research highlights the severe implications of vulnerabilities in these often-overlooked lower layers, which lack cryptographic protections and can lead to remote code execution or information leakage, even after authentication and key agreement.

Watch on YouTube · Slides

Visual summary for LLFuzz: An Over-the-Air Dynamic Testing Framework for Cellular Baseband Lower Layers by Tuan Dinh Hoang
Visual summary for LLFuzz: An Over-the-Air Dynamic Testing Framework for Cellular Baseband Lower Layers by Tuan Dinh Hoang

Key moments

  1. 0:00 Introducing LLFuzz: Over-the-air testing for baseband lower layers
  2. 1:50 Why existing baseband testing methods fall short
  3. 3:00 Understanding the intricate structures of lower layers
  4. 4:30 Overcoming complex packet structures with specification guidance
  5. 5:20 Channel-driven stateful testing for multiple logical channels
  6. 6:30 Configuration-aware testing for varying packet structures
  7. 7:20 Overview of the LLFuzz architecture and components
  8. 8:20 LLFuzz's test case generation and crash detection methods

LLFuzz: An Over-the-Air Dynamic Testing Framework for Cellular Baseband Lower Layers

Speakers: Tuan Dinh Hoang

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=5qg9uJXzdbY

Overview

This talk introduces LLFuzz, an innovative over-the-air dynamic testing framework designed to uncover memory corruption vulnerabilities within the lower layers of cellular basebands. Presented by Tuan Dinh Hoang, a PhD student from CEC Lab, KAIST Korea, LLFuzz addresses a critical gap in existing security research, which has predominantly focused on the higher layers (Layer 3) of cellular protocol stacks. The research highlights the severe implications of vulnerabilities in these often-overlooked lower layers, which lack cryptographic protections and can lead to remote code execution or information leakage, even after authentication and key agreement.

The significance of LLFuzz lies in its systematic approach to testing modern lower layers (LTE and 5G) across various vendors without requiring internal firmware access. By leveraging an over-the-air (OTA) testing methodology, LLFuzz is able to perform stateful testing, which is crucial for the complex, channel-driven nature of lower-layer protocols. The framework's ability to identify previously unknown vulnerabilities in widely deployed commercial basebands underscores the urgent need for enhanced security scrutiny in this fundamental component of mobile communication.

Background

▶ Watch: Introducing LLFuzz: Over-the-air testing for baseband lower layers (0:00)

Cellular networks rely on three primary components: User Equipment (UE), base stations, and the core network. Within the UE, the baseband processor (BP) is a critical component responsible for all wireless communication between the UE and the base station. Its protocol stack is logically divided into three main layers: Layer 1 (Physical Layer), Layer 2 (comprising MAC, RLC, and PDCP sub-layers), and Layer 3 (including protocols like NAS and RRC). While the baseband processor is fundamental to cellular operation, it is a frequent target for attackers due to its critical role and the complexity of its codebase, often written in C and C++.

Vulnerabilities in basebands typically fall into two categories: specification vulnerabilities (flexification) and implementation vulnerabilities, which include protocol-level and memory corruption issues. Memory corruption is particularly prevalent because basebands handle numerous functions to decode downlink packets from base stations, making them susceptible to errors when processing malformed data. Exploiting these vulnerabilities can lead to severe consequences, such as remote code execution (RCE) or information leakage, making them a significant focus for both academia and industry.

Despite the critical nature of baseband security, most prior research and testing efforts have concentrated on Layer 3. For instance, tools like BasePack reverse-engineer Layer 3 functions for specification comparison, while others like FirmWire emulate Samsung and MediaTek baseband Layer 3 functions using fuzzing tools like AFL. More recently, Professor Sites' group introduced LowRISC, an emulator supporting Layer 3 of 5G basebands. In contrast, Layer 2 has received considerably less attention. Good extended FirmWire for Layer 2 of GPRS, and 5G-Fuzz utilized over-the-air fuzzing with Layer 2 support. However, these efforts did not cover modern LTE and 5G lower layers comprehensively, and 5G-Fuzz primarily focused on pre-authentication states with random mutation, which is often inefficient. Crucially, state-of-the-art emulators largely lack support for LTE and 5G lower layers, leaving a significant attack surface unexplored.

The lower layers present unique challenges that complicate systematic testing. First, they possess diverse and complex packet structures. The Physical Layer decodes downlink control information (DCI) before data packets, while the MAC layer has structures for Random Access Response and MAC Control Elements. RLC and PDCP layers use dedicated control packets for transmission status reporting. Moreover, these layers utilize multiple logical channels, each mapping different data types to specific packet structures. This means random packet generation is ineffective; testers must consider channel mappings. Second, these logical channels are not always active but are enabled by specific signaling messages during connection establishment. For example, during the initial connection phase, only Random Access Response messages are processed in the MAC layer, followed by the Common Control Channel (CCCH) for signaling. Sending test cases to inactive channels results in rejection. Third, even the same logical channel can have varying packet structures depending on its configuration, which is set via RRC connection reconfiguration messages. For instance, an RLC packet's sequence number field can be 5-bit or 10-bit based on RRC configuration. These complexities necessitate a systematic, state-aware, and configuration-aware approach, which LLFuzz provides.

Key Findings

▶ Watch: Understanding the intricate structures of lower layers (3:00)

LLFuzz successfully identified a total of 11 previously unknown memory corruption vulnerabilities across 15 commercial basebands from five major vendors: Qualcomm, MediaTek, Exynos, Google Tensor, and Huawei. These findings represent a significant contribution to cellular security, revealing critical weaknesses in widely deployed components.

Specifically, the framework uncovered:

  • 9 memory corruptions in LTE basebands: These were distributed across Layer 2 sub-layers, with two in PDCP, two in RLC, and five in the MAC layer. The distribution highlights that all sub-layers of Layer 2 are susceptible to such vulnerabilities.
  • 2 additional bugs in 5G basebands: One was found in the PDCP layer and another in the RLC layer during a focused 5G prototype development. These findings demonstrate LLFuzz's applicability to the latest cellular generation, even with limited development time.

The identified vulnerabilities have been assigned 9 CVEs to date, with many affecting over 80 commercial basebands. This widespread impact is particularly concerning given that these basebands are integrated into a vast array of devices, including smartphones, IoT devices, and even cars. The talk emphasizes that the bugs found in the lower layers are more critical because they remain exploitable even after successful authentication and key agreement. This is due to the inherent design of lower layers, which are not cryptographically protected, making them vulnerable to attack models such as "seek over" or man-in-the-middle attacks.

A crucial lesson learned from the disclosure process is the significant supply chain problem. Baseband vendors indicated that even after releasing a fix, there is no guarantee that it will be patched on all affected devices. The application of patches often depends on the device vendors, leaving many products potentially unpatched and vulnerable for extended periods. This issue exacerbates the risk posed by these lower-layer vulnerabilities, as a large installed base of devices could remain exposed.

Technical Deep Dive

▶ Watch: Channel-driven stateful testing for multiple logical channels (5:20)

LLFuzz is designed as a systematic, over-the-air dynamic testing framework to detect memory corruptions in the lower layers of cellular basebands (Physical, MAC, RLC, PDCP). It is built on top of an open-source full-stack 4G base station, specifically srsRAN (formerly srsLTE), and is implemented in C and C++ with over 11,000 lines of code. The framework comprises three main components: Specification Analysis, Over-the-Air Testing, and Post Analysis.

LLFuzz addresses the three core challenges of lower-layer testing through specific design choices:

  1. Challenge 1: Complex Packet Structures in Over-the-Air Testing: Lower layers have numerous packet structures and fields, but only a small subset is used in practice. Random mutation is inefficient due to low OTA testing speed and early rejection by basebands, and commercial basebands are black boxes, preventing coverage-guided fuzzing.
  • LLFuzz's Solution: Specification-Guided Test Generation: LLFuzz carefully reviews 3GPP specifications to identify 19 structures in the MAC layer, 18 in RLC, 17 in PDCP, and 11 DCI structures in the Physical layer. This detailed analysis allows the generation of diverse, standard-compliant packets, which increases the likelihood of reaching vulnerable code paths and aids in root cause analysis.

The test case generation process involves four steps:

  1. Pick a specific structure and generate a legitimate packet with various components.
  2. Mutate the header, focusing on boundary and reserved values.
  3. Mutate only the payload portions that are processed by the target testing layer.
  4. Map the generated test cases to corresponding logical channels by specifying the logical channel ID and the length field in the MAC headers.
  1. Challenge 2: Diverse Packet Structures Across Multiple Logical Channels: Logical channels dictate the packet structure, and they are not always active, being enabled by specific signaling messages during connection establishment. Sending tests to inactive channels leads to rejection.
  • LLFuzz's Solution: Channel-Driven Stateful Testing: LLFuzz defines four general-oriented states based on the establishment of logical channels. This approach ensures that LLFuzz knows exactly which channels are active at any given point and which packet structures are appropriate for generating and sending test cases. For example, in the initial phase, only Random Access Response messages are processed; subsequently, the CCCH channel becomes active for signaling, requiring different MAC layer packet structures.
  1. Challenge 3: Configuration-Dependent Packet Structures: Even the same logical channel can have different packet structures based on how it is configured during connection establishment (e.g., RLC sequence number field can be 5-bit or 10-bit based on RRC configuration messages).
  • LLFuzz's Solution: Configuration-Aware Testing: LLFuzz actively modifies RRC reconfiguration messages to deliver target configurations to the baseband. After configuring the UE, it generates and sends test cases that respect the newly applied configuration, ensuring comprehensive coverage of different operational modes.

The Over-the-Air Testing component orchestrates the fuzzing process. It first triggers an attach procedure by sending ping messages or switching the app mode via the ADB interface. Then, it modifies RRC reconfiguration messages to set the target configuration for the baseband. Once the baseband reaches the desired testing state (determined by the channel-driven stateful approach), LLFuzz sends the corresponding test cases. During and after this, it monitors the ADB logcat for crash indicators.

Demo / Proof of Concept

▶ Watch: Configuration-aware testing for varying packet structures (6:30)

The talk describes the practical implementation of LLFuzz's testing procedure, which serves as its proof-of-concept for discovering vulnerabilities. The process is fully automated for initial crash detection, followed by a semi-manual validation phase.

The automated testing procedure works as follows:

  1. Trigger Attach Procedure: The process begins by initiating a connection between the UE and the LLFuzz-controlled base station. This is achieved by sending ping messages or programmatically switching the application mode on the UE via the Android Debug Bridge (ADB) interface.
  2. Deliver Target Configuration: LLFuzz actively intercepts and modifies RRC reconfiguration messages. This allows it to deliver specific, desired configurations to the baseband under test, ensuring that different operational modes and packet structure variations are explored.
  3. Send Test Cases: Once the baseband is in the target testing state (as determined by the channel-driven stateful logic), LLFuzz sends the carefully crafted, specification-guided test cases over the air. These test cases are designed to probe various fields and structures within the MAC, RLC, PDCP, and Physical layers.
  4. Monitor for Crashes: Immediately after sending test cases, LLFuzz continuously monitors the ADB logcat output from the connected UE. The framework identifies potential crashes by looking for specific strings that commonly appear when a baseband experiences a critical error, such as "radio off," "unavailable model reset," or "everybody panic."
  5. Log and Reset: If a crash indicator is detected, LLFuzz logs the "crash candidate" to its database. If no crash is detected, or after a crash is logged, LLFuzz disconnects the UE to reset the baseband and initiate a new testing session, ensuring isolation between test cases.

While the ADB logcat provides an efficient automated method for initial crash detection, it can sometimes produce false positives. To address this, a post-analysis phase is employed:

  • Vendor Debug Mode Validation: For crash candidates identified during automated testing, the phone is switched to a vendor-specific debug mode. In this mode, if the baseband truly crashes, the phone will typically display a black screen, which is a more reliable and definitive indicator of a baseband crash compared to logcat messages alone.
  • Manual Restart: A limitation of this highly reliable validation method is that after each confirmed crash in vendor debug mode, the phone must be manually restarted. This makes it unsuitable for full automation but is essential for verifying the authenticity of discovered vulnerabilities.

This systematic approach, combining automated fuzzing with rigorous manual validation, allowed LLFuzz to effectively pinpoint 11 previously unknown vulnerabilities across various commercial basebands.

Defensive Implications

▶ Watch: LLFuzz's test case generation and crash detection methods (8:20)

The findings from LLFuzz carry significant implications for the security of cellular networks and the devices that rely on them. The discovery of numerous memory corruption vulnerabilities in the lower layers of basebands highlights several critical areas where defensive strategies must be strengthened:

  1. Urgent Need for Lower-Layer Security Audits: The talk unequivocally states that memory corruptions are still common in lower layers, and their specifications often lack specific security testing requirements for these types of bugs. This indicates a systemic oversight that needs to be corrected. Defenders, including baseband vendors and device manufacturers, must prioritize comprehensive security audits and fuzzing efforts specifically targeting Layers 1 and 2, mimicking LLFuzz's stateful and configuration-aware approach.
  2. Criticality of Lower-Layer Exploits: Unlike higher-layer vulnerabilities that might be mitigated by cryptographic protections, lower-layer bugs are particularly dangerous because they are not cryptographically protected by design. This means they remain exploitable even after successful authentication and key agreement. Attackers can leverage models like "seek over" (where a malicious base station tricks a UE into connecting) or man-in-the-middle attacks to exploit these vulnerabilities, potentially leading to remote code execution or denial of service on a victim's device. This necessitates a shift in threat modeling to account for post-authentication, lower-layer attacks.
  3. Addressing the Supply Chain Problem: A major defensive challenge identified is the cellular supply chain. Even when baseband vendors release fixes for vulnerabilities, there is "no warranty that it will be patched on all devices." The responsibility for applying fixes often falls to device vendors, leading to a fragmented patching ecosystem where many products can remain unpatched indefinitely. This "supply chain problem" requires multi-stakeholder collaboration, potentially involving regulatory pressure, industry standards, or coordinated disclosure programs to ensure that critical security updates consistently reach end-users. Consumers and enterprises should also be educated on the importance of timely updates and choose devices from vendors with strong patch delivery track records.
  4. Uplink Security Concerns: The research focused on downlink vulnerabilities, but the speaker raises a critical point about uplink. A single memory corruption in the uplink processing of a base station or core network could have even more severe consequences, potentially taking down an entire base station or a significant portion of the core network. This highlights an area for future research and defensive focus, as the impact of such an attack could be catastrophic for network availability and integrity.
  5. Enhanced Monitoring and Anomaly Detection: While baseband crashes are often detected via logcat messages or debug modes, defenders should explore more sophisticated anomaly detection systems at the network edge and within UEs. These systems could identify unusual packet structures, abnormal device behavior, or unexpected baseband resets that might indicate an ongoing attack or the presence of an unpatched vulnerability.

In summary, LLFuzz’s findings serve as a stark warning that the foundational layers of cellular communication are ripe for exploitation. A robust defense requires not only technical solutions like thorough fuzzing but also systemic changes in how the industry approaches security, patching, and supply chain responsibility.

Key Takeaways

  • Memory corruptions are prevalent in lower layers: Despite their critical role, the MAC, RLC, PDCP, and Physical layers of cellular basebands still contain numerous memory corruption vulnerabilities, often overlooked in security testing.
  • Lower-layer bugs are highly critical: These vulnerabilities lack cryptographic protection by design, making them exploitable even after authentication and key agreement, and susceptible to remote attacks via malicious base stations or man-in-the-middle scenarios.
  • LLFuzz provides a systematic OTA testing approach: The framework effectively addresses the complexities of lower-layer protocols through specification-guided test generation, channel-driven stateful testing, and configuration-aware testing.
  • Significant real-world impact: LLFuzz discovered 11 previously unknown bugs across 15 commercial basebands from five major vendors, leading to 9 CVEs and affecting over 80 commercial products, including smartphones, IoT devices, and cars.
  • Supply chain issues hinder patching: Even when fixes are released by baseband vendors, the fragmented supply chain often prevents these patches from reaching all affected devices, leaving a vast number of products vulnerable.
  • Uplink security is a future concern: While LLFuzz focused on downlink, memory corruptions in uplink processing could critically impact base stations or the core network, highlighting an area for future research and defensive efforts.

About the Speaker(s)

Tuan Dinh Hoang is a PhD student from CEC Lab, KAIST Korea. His research focuses on cellular security, particularly the identification of vulnerabilities in baseband processors. The work presented on LLFuzz is a collaborative effort with Professor Jojun, Professor Insu, and Professor Yang, reflecting a deep academic engagement in addressing critical security challenges within wireless communication systems.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid original research from a PhD student who clearly did the actual work — 11 previously unknown memory corruptions, 9 CVEs, five major vendors, real OTA framework with 11K+ lines of code. The gap being addressed (L2 baseband fuzzing for LTE/5G) is real and underserved, and the tri-part solution to stateful OTA testing is technically grounded.

Heather Calloway (CISO) — WEAK

Technically credible work — 11 CVEs across five major vendors is not nothing — but this is a research paper delivered as a conference talk, and it never bridges to the people who need to act on it. The supply chain disclosure at the end is the most operationally important finding in the entire presentation, and it gets a paragraph.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)