GPUHammer: Rowhammer Attacks on GPU Memories are Practical

Chris S. Lin (University of Toronto)

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Hardware Security 3: Side-Channel and Fault Injection Attacks

Overview

The talk "GPUHammer: Rowhammer Attacks on GPU Memories are Practical" presents a groundbreaking study demonstrating the first practical Rowhammer attacks specifically targeting Graphics Processing Unit (GPU) memories. Presented by Chris Joyce, a researcher from the University of Toronto, this work highlights a critical security vulnerability in a hardware component increasingly central to modern computing, particularly in machine learning (ML) and high-performance computing. The research underscores that GPUs, much like CPUs, are susceptible to Rowhammer-induced bit flips, posing a significant threat to data integrity and the security of sensitive applications.

Watch on YouTube · Read the paper · Download the PDF (PDF) · Slides

Paper abstract

eSIM (Embedded Subscriber Identity Module) technology is rapidly reshaping mobile connectivity by enabling users to activate cellular services without a physical SIM card. While the flexibility of remote provisioning improves convenience and scalability, particularly for international travelers, it also introduces complex and underexplored privacy and security risks. This paper presents an empirical investigation of how eSIM adoption affects user privacy, focusing on routing transparency, reseller access, and profile control. We first show how travel eSIMs often route user data through third-party networks, including Chinese infrastructure, regardless of user location. This raises concerns about jurisdictional exposure. Second, we analyze the implications of opaque provisioning workflows, documenting how resellers can access sensitive user data, proactively communicate with devices, and assign public IPs without user awareness. Third, we validate operational risks such as deletion failures and profile lock-in using a private LTE testbed. In addition to these empirical contributions, we reflect on the evolving threat landscape of eSIM technology and analyze the shifting trust boundaries introduced by its global provisioning architecture. We conclude with actionable recommendations for improving eSIM transparency, user control, and regulatory enforcement as the technology becomes widespread across smartphones, IoT deployments, and private networks.

Visual summary for GPUHammer: Rowhammer Attacks on GPU Memories are Practical by Chris S. Lin
Visual summary for GPUHammer: Rowhammer Attacks on GPU Memories are Practical by Chris S. Lin

Key moments

  1. 0:00 Introduction: GPUHammer and GPU vulnerability
  2. 1:18 Basic Rowhammer workflow on GPUs
  3. 2:18 Overview of three main GPU Rowhammer challenges
  4. 2:48 Challenge 1: Virtual memory mapping via timing
  5. 4:40 Challenge 2: Overcoming high GPU memory latency
  6. 6:18 Solution: Multi-warp hammering for high intensity
  7. 7:00 Challenge 3: Bypassing DRAM defenses like TRR

GPUHammer: Rowhammer Attacks on GPU Memories are Practical

Speakers: Chris S. Lin

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=KpRmVMTS3b0

Overview

The talk "GPUHammer: Rowhammer Attacks on GPU Memories are Practical" presents a groundbreaking study demonstrating the first practical Rowhammer attacks specifically targeting Graphics Processing Unit (GPU) memories. Presented by Chris Joyce, a researcher from the University of Toronto, this work highlights a critical security vulnerability in a hardware component increasingly central to modern computing, particularly in machine learning (ML) and high-performance computing. The research underscores that GPUs, much like CPUs, are susceptible to Rowhammer-induced bit flips, posing a significant threat to data integrity and the security of sensitive applications.

This research is particularly pertinent given the widespread adoption of GPUs in cloud environments and shared computing infrastructures, where multiple users or applications might share the same physical hardware. The ability to induce bit flips on GPU memory opens avenues for various malicious exploits, ranging from data corruption to more sophisticated attacks like privilege escalation or the subversion of ML models. The talk meticulously details the challenges in adapting Rowhammer to the unique architecture and memory characteristics of GPUs and presents novel techniques to overcome these hurdles, ultimately proving the practicality of such attacks.

The implications of GPUHammer are far-reaching. As ML models become more complex and critical, their integrity is paramount. A successful Rowhammer attack could lead to manipulated model outputs, backdoors, or complete model incapacitation, with severe consequences for AI-driven systems. The work serves as a stark reminder that hardware vulnerabilities persist across different computing paradigms and necessitates a re-evaluation of security postures in GPU-accelerated environments.

Background

▶ Watch: Introduction: GPUHammer and GPU vulnerability (0:00)

The Rowhammer vulnerability, first discovered in 2014, exposed a critical flaw in Dynamic Random Access Memory (DRAM) chips. DRAM cells are organized into banks of rows, and rapid, repeated accesses (or "activations") to a specific "aggressor" row can disturb the electrical charge in physically adjacent "victim" rows. This disturbance can lead to bit flips, where a '0' changes to a '1' or vice versa, compromising data integrity.

On Central Processing Units (CPUs), Rowhammer exploits are well-documented across both DDR and LPDDR memory types. These exploits have demonstrated severe consequences, including privilege escalation through instruction modification, page table tampering, and the extraction of confidential data such as cryptographic keys. However, the focus has historically been on CPU-based DRAMs.

The security landscape for GPUs, despite their increasing computational power and role in processing sensitive data (e.g., ML models, financial simulations), has largely been underexplored regarding Rowhammer. GPUs employ their own dedicated DRAM modules, such as GDDR6 or HBM2E, which share the fundamental architecture susceptible to Rowhammer. The threat model for GPUs often involves time-shared scenarios, where different users or processes run their code at separate intervals but share the same physical DRAM. This collocation of data makes GPUs inherently vulnerable to Rowhammer attacks, where a malicious actor could run code that hammers memory while a victim's sensitive data resides in an adjacent, vulnerable row.

To launch a Rowhammer attack on any DRAM, the basic workflow involves:

  1. Uncached Memory Access: The attacker must access memory addresses that are not served from a cache, ensuring the request goes directly to the DRAM.
  2. Row Activation: When data from a specific physical row (e.g., row 1) is requested, an ACT (activation) command is issued to that row. This command opens the row and brings its data into the row buffer for read/write operations. These ACT commands are the primary cause of Rowhammer disturbances.
  3. Row Buffer Eviction: The row buffer acts as a cache. Subsequent accesses to the same row are served directly from the buffer without needing another ACT command. To continuously trigger ACT commands, the attacker must force the eviction of the currently open row from the buffer by activating a different row. By repeatedly activating two or more aggressor rows, the attacker can rapidly trigger ACT commands, leading to bit flips in neighboring victim rows.

While the fundamental concept remains, several significant challenges arise when attempting to port Rowhammer attacks from CPUs to GPUs:

  1. Virtual Memory Mapping: GPUs operate with virtual memory addresses, making it difficult to determine the physical bank and row an address maps to. This knowledge is crucial for selecting aggressor rows that are physically adjacent to target victim rows.
  2. High Memory Latency: GPUs exhibit significantly higher memory latency compared to CPUs (typically 4-5 times higher). This makes it challenging to achieve the high activation rates necessary to induce bit flips within the strict refresh windows of DRAM.
  3. Modern DRAM Defenses: Contemporary DRAM chips incorporate built-in mitigations, such as Target Row Refresh (TRR), designed to counteract Rowhammer. These defenses track recently accessed rows and periodically refresh them to prevent charge leakage, requiring attackers to develop bypass techniques. Overcoming these GPU-specific challenges is central to GPUHammer's practicality.

Key Findings

▶ Watch: Overview of three main GPU Rowhammer challenges (2:18)

The GPUHammer research successfully demonstrates the first practical Rowhammer attacks on modern GPUs, revealing several critical findings:

  1. Practical Bit Flips on GDDR6: The researchers successfully induced bit flips on an NVIDIA RTX A6000 GPU utilizing GDDR6 memory. This confirms that modern GPU memory architectures are indeed vulnerable to Rowhammer, despite their distinct characteristics from CPU DRAMs. While other GPUs tested (A100 with HBM2E, RTX 3080 with GDDR6) did not exhibit bit flips, this highlights potential differences in memory chip resilience, higher Rowhammer thresholds, or unaddressed in-DRAM mitigations across different GPU models.
  2. Novel Techniques for GPU Rowhammer: The project introduces a suite of innovative techniques tailored to overcome the unique challenges of GPU architectures:
  • An effective timing side-channel scheme to reverse-engineer virtual-to-physical memory mappings for GPU DRAM.
  • Multi-warp hammering strategies to achieve unprecedented activation rates (up to 93% of theoretical maximum), bypassing the high memory latency of GPUs.
  • An extended multi-thread synchronization mechanism to bypass in-DRAM defenses like Target Row Refresh (TRR) by consistently overflowing and manipulating its internal tracker.
  1. Characterization of GPU Bit Flips: The study provides detailed characterizations of the observed bit flips, identifying that:
  • Bit flips were primarily triggered by aggressor rows on one side only of the victim row.
  • Rows at r+/-2 (two rows away from the victim) were the most effective aggressors, suggesting a non-contiguous physical row layout.
  • The Rowhammer threshold on the tested GDDR6 memory was found to be between 12,300 and 16,000 activations per single DRAM row.
  • The TRR tracker size on the A6000 was characterized, requiring at least 17 aggressor rows to overflow it, suggesting an effective tracker capacity of 16 entries per bank.
  1. Demonstrated ML Model Exploitation: The researchers successfully demonstrated an accuracy degradation attack on popular Deep Neural Network (DNN) ML models. By strategically flipping a single bit in the floating-point exponent of random model weights, they could reduce model accuracy from 80% to 0% in under 10 attempts. This highlights a severe practical consequence of GPU Rowhammer for AI systems.
  2. Responsible Disclosure and Mitigation: The vulnerability was responsibly disclosed to NVIDIA, who confirmed the issue and recommended enabling ECC (Error-Correcting Code) for affected GPUs as a mitigation. However, the research noted that enabling ECC resulted in an observed up to 10% slowdown on their GPU, emphasizing the performance trade-offs of current hardware mitigations. The authors advocate for more fundamental hardware-level solutions in future DRAM designs.

These findings collectively establish GPUHammer as a significant contribution to hardware security, exposing a critical vulnerability in a widely used computing platform and offering a foundation for future research into GPU-specific hardware attacks and defenses.

Technical Deep Dive

▶ Watch: Challenge 1: Virtual memory mapping via timing (2:48)

The GPUHammer attack meticulously addresses three primary technical challenges inherent in adapting Rowhammer to GPU architectures.

Challenge 1: Virtual Memory Mapping

The first hurdle is determining the physical DRAM bank and row corresponding to a given virtual memory address. Unlike CPUs, GPU memory controllers often employ complex, undocumented mapping functions. To overcome this, GPUHammer leverages timing side-channels, a technique previously used in CPU Rowhammer research (e.g., "Drama").

The core idea is to measure the access latency between pairs of memory addresses:

  • If two addresses, 0xA and 0xB, reside in different DRAM banks, they can be accessed in parallel. Each bank has its own row buffer, leading to low access latency.
  • If 0xA and 0xB are in the same DRAM bank, they must contend for the single row buffer within that bank. This forces them to take turns, resulting in higher access latency.

The latency measurement technique involves:

  1. Using GPU threads to access an address pair (0xA and 0xB).
  2. Clearing the cache for both addresses to ensure requests reach DRAM.
  3. Ensuring threads proceed only after both accesses have completed, capturing the slowest of the two latencies.
  4. Taking the minimum of 10 measurements to mitigate system noise and transient spikes.

Initial visualization of access time differences across the entire memory layout revealed an concerning overlap around 370 nanoseconds, making same-bank and different-bank addresses indistinguishable. This phenomenon is attributed to the Non-Uniform Memory Access (NUMA) effect, where memory locations physically distant from the GPU can exhibit longer latencies even without contention. To filter out NUMA effects, the researchers applied an insight: addresses within the same bank must reside on the same physical DRAM chip, implying physical proximity. They filtered out addresses that showed high access latency even when accessed individually, indicating they were "too far apart" to be in the same bank. After this NUMA filtering, a clear cut-off emerged, enabling reliable identification of addresses belonging to the same bank and, subsequently, the same row by observing the absence of such access conflicts.

Challenge 2: High Memory Latency and Hammering Intensity

GPUs are notorious for their high memory latency, which is typically 4-5 times higher than that of CPUs. This poses a significant challenge for Rowhammer, which relies on rapid activations to induce bit flips before DRAM cells are refreshed. DRAM cells periodically leak charge and must be refreshed within a TRFW (Refresh Window), typically 32 milliseconds. Furthermore, individual refresh commands are issued every TRFE (Refresh Interval), around 1.9 microseconds on GDDR6. The goal is to maximize activations within the TRFW.

A naive single-thread hammering loop, similar to CPU approaches, proved highly suboptimal. The thread issues a load request, triggering a row activation. However, a significant portion of time is then spent on data transfer back to the Streaming Multiprocessor (SM), during which the DRAM remains idle. This leads to only about 15% of the theoretical maximum hammering intensity.

To improve intensity, the researchers explored multi-thread hammering. By using multiple threads, memory requests can be issued in parallel, overlapping delays and increasing throughput. This significantly boosted intensity to 80% of the theoretical maximum. However, a caveat exists: threads within a warp (typically 32 threads) execute in lock-step. Even if one thread receives its data early, it must wait for the slowest thread in the warp to complete before proceeding, introducing idle time.

The most effective solution developed is multi-warp hammering. In this approach, one effective thread per warp is utilized, and multiple warps are employed to issue requests. Crucially, warps are scheduled independently. This means one warp can send a subsequent memory request as soon as its previous request completes, without waiting for other warps. This strategy further minimizes idle time on the DRAM, achieving up to 620,000 activations per refresh window with eight or more warps, reaching an impressive 93% of the theoretical maximum. This intensity proved sufficient to reliably trigger bit flips.

Challenge 3: Bypassing In-DRAM Defenses (Target Row Refresh - TRR)

Modern DRAM chips incorporate defenses like Target Row Refresh (TRR) to mitigate Rowhammer. TRR mechanisms typically maintain a fixed-size tracker that records recently accessed rows. If a row is accessed frequently enough, TRR will proactively refresh it, preventing bit flips.

Prior work on CPU Rowhammer proposed two main techniques to bypass TRR:

  1. Trespass: This involves many-sided hammering, activating more distinct aggressor rows than the tracker can hold. When the tracker is full, activating a new aggressor forces an older entry out, which then escapes refreshment. For example, a 4-entry tracker would evict an entry if five distinct aggressors are hammered.
  2. Smash: This technique synchronizes hammering with mitigated refreshes. By deliberately creating pauses or "holes" in the hammering pattern, the attacker ensures that TRR refreshes occur at predictable intervals, consistently targeting and evicting the same entry from the tracker.

GPUHammer extends this synchronization to a multi-thread setting. The researchers insert the same delay after each warp's hammering rounds, creating aligned gaps that allow TRR's mitigated refreshes to occur predictably. To identify the optimal delay for synchronization, another timing trick is employed. After each round of hammering, additional operations (e.g., additions) are introduced to create controlled delays. If this additional delay overlaps with the TRR refresh latency, the total time per hammering round will plateau instead of increasing linearly. Observing this plateau (e.g., around 56 additions in the presented data) indicates that the hammering is aligned with the DRAM refresh periods. By aligning hammering with DRAM refresh periods, the attack can reliably manipulate the TRR tracker, ensuring that specific victim rows are not refreshed and thus remain vulnerable to bit flips.

Demo / Proof of Concept

▶ Watch: Solution: Multi-warp hammering for high intensity (6:18)

With all three technical challenges addressed, the GPUHammer team proceeded to a practical demonstration and characterization of bit flips on real GPUs.

Bit Flip Campaign and Results

A comprehensive hammering campaign was conducted on three distinct GPU models:

  • An RTX A6000 with GDDR6 memory.
  • An A100 with HBM2E memory.
  • An RTX 3080 also with GDDR6 memory.

For each GPU, four memory banks were hammered using 8 to 24 aggressor rows and a checkered data pattern. This extensive testing involved approximately 30 hours per bank per GPU. The results showed clear success on the RTX A6000, where eight bit flips were observed across all tested banks. Crucially, no bit flips were observed on the A100 or RTX 3080. The researchers hypothesize that this could be due to differences in memory chip design, higher inherent Rowhammer thresholds in those specific DRAM modules, or unaddressed in-DRAM mitigations that were more effective on these particular GPUs.

Characterization of Bit Flips

Upon observing bit flips on the A6000, further characterization was performed:

  • Critical Aggressor Rows: The study identified the specific aggressor rows responsible for each bit flip. It was observed that bit flips could only be triggered by aggressor rows from one side of the victim row, and rows at r+/-2 (two rows away from the victim) were the most effective. This suggests a physical memory layout that is not strictly contiguous or symmetrical.
  • Rowhammer Threshold: To determine the minimum intensity required, the researchers characterized the Rowhammer threshold on GDDR6. They found that it required at least 12,300 activations to a single DRAM row before observing a bit flip, with some flips requiring up to 16,000 activations.
  • TRR Tracker Size: By observing the number of aggressor rows needed to trigger bit flips, the team characterized the TRR tracker size on the A6000. Bit flips only occurred when the number of aggressor rows was 17 or more. This implies that the TRR tracker can hold 16 rows per bank, as 17 aggressors are needed to overflow this capacity and force an entry out.

ML Model Exploit: Accuracy Degradation

The most compelling proof of concept involved demonstrating a practical exploit against Machine Learning models. The threat model assumed a time-slice scenario, where a victim and an attacker run their code at separate time intervals but share the same underlying GPU DRAM.

The attack requires placing victim data (specifically, ML model weights) into a vulnerable row. To achieve this, the attacker first allocates the entire GPU memory and then strategically frees "holes" in the identified vulnerable regions. During the victim's time slice, when the ML model is loaded, its data is forced into these pre-prepared holes, making it susceptible to bit flips. In the subsequent time slice, the attacker executes the multi-warp hammering process.

The target of the exploit was the accuracy degradation of DNNs. Prior work (e.g., "Terminal Brain Damage") has shown that a single bit flip in the floating-point exponent of a model weight can drastically alter the weight's value, severely degrading model accuracy. By flipping the top bit in the exponent of random weights within five popular DNN ML models, the researchers demonstrated that in under 10 attempts, they could degrade model accuracy from approximately 80% to 0%. This highlights the devastating impact Rowhammer can have on the reliability and trustworthiness of AI systems.

Defensive Implications

▶ Watch: Challenge 3: Bypassing DRAM defenses like TRR (7:00)

The GPUHammer research carries significant implications for defenders and highlights the need for robust security measures in GPU-accelerated environments.

Upon responsible disclosure, NVIDIA confirmed the vulnerability and issued a security notice. Their primary recommendation for mitigation is to enable Error-Correcting Code (ECC) on GPUs where available. ECC memory can detect and correct single-bit errors and detect some multi-bit errors, thereby counteracting Rowhammer-induced bit flips. However, the researchers observed a notable drawback: enabling ECC on their tested GPU resulted in an up to 10% slowdown in performance. This performance overhead presents a difficult trade-off for users and organizations, especially in scenarios where maximum computational throughput is critical, such as in large-scale ML training or high-performance computing.

The core message from the GPUHammer team is that Rowhammer is fundamentally a hardware flaw. While ECC offers a software-level (or rather, memory controller-level) mitigation, it does not address the root cause in the DRAM cells themselves. Therefore, the researchers strongly urge for the implementation of principal mitigations in future DRAM designs. This could involve more robust in-DRAM defenses that are less susceptible to bypass techniques, or entirely new memory architectures that inherently prevent charge leakage and inter-row interference.

For immediate defensive actions, organizations operating GPU-accelerated systems, particularly those handling sensitive data, critical ML models, or operating in multi-tenant cloud environments, should:

  • Prioritize ECC: Enable ECC on GPUs where data integrity is paramount, accepting the potential performance overhead. This is crucial for applications like financial modeling, medical imaging, or any system where a single bit flip could have catastrophic consequences.
  • Isolate Workloads: In shared GPU environments, implement stronger workload isolation mechanisms to prevent malicious co-tenants from launching Rowhammer attacks. While perfect isolation at the hardware level is challenging, strategies like dedicated physical GPUs per tenant or robust hypervisor-level memory management could help.
  • Monitor and Audit: Implement monitoring solutions that can detect unusual memory access patterns or unexpected data corruption, which could indicate a Rowhammer attack.
  • Stay Informed: Keep abreast of new research and vendor advisories regarding hardware vulnerabilities and recommended mitigations.

The GPUHammer research underscores that hardware security must be a continuous focus, not just for CPUs but for all critical processing units. The increasing reliance on GPUs for sensitive tasks necessitates a proactive approach to addressing such fundamental hardware vulnerabilities.

Key Takeaways

  • GPUHammer is the first practical Rowhammer attack on modern GPUs, specifically demonstrated on an NVIDIA RTX A6000 with GDDR6 memory.
  • The research introduces novel techniques to overcome GPU-specific challenges, including timing side-channels for memory mapping, multi-warp hammering for high activation intensity, and multi-thread synchronization to bypass Target Row Refresh (TRR) defenses.
  • Bit flips were characterized, showing that r+/-2 aggressor rows were most effective and occurred from one side only, with a Rowhammer threshold of 12,300 to 16,000 activations and a TRR tracker capacity of 16 entries per bank.
  • A practical exploit was demonstrated, achieving accuracy degradation in DNN ML models from 80% to 0% in under 10 attempts by flipping a single bit in a floating-point exponent.
  • NVIDIA confirmed the vulnerability and recommends enabling ECC, but this comes with an observed up to 10% performance slowdown.
  • The research emphasizes that Rowhammer is a fundamental hardware flaw that necessitates principal mitigations in future DRAM designs, beyond current software-level patches.

About the Speaker(s)

The research presented in "GPUHammer: Rowhammer Attacks on GPU Memories are Practical" was conducted by Chris S. Lin, a researcher from the University of Toronto. The talk itself was delivered by Chris Joyce. This work was supervised by Professor Gurush at the University of Toronto, highlighting expertise in computer architecture and security research.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid, well-scoped hardware security research that extends Rowhammer — a mature but still-evolving attack class — into genuinely underexplored territory: GPU DRAM. The three-challenge framework (address mapping, hammering intensity, TRR bypass) is methodically solved with novel adaptations, and the ML accuracy-degradation demo gives the work real-world bite beyond a pure academic exercise.

Heather Calloway (CISO) — WEAK

Technically rigorous work that establishes GPU Rowhammer as real — but the paper does the heavy lifting, not the talk. The governance and operational gaps are significant enough that this doesn't land where it needs to for anyone running a cloud security program or advising a board on AI infrastructure risk.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)