ECC.fail: Mounting Rowhammer Attacks on DDR4 Servers with ECC Memory

Nureddin Kamadan

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Hardware Security 3: Side-Channel and Fault Injection Attacks

Overview

The "ECC.fail" talk presented at USENIX Security unveils a groundbreaking Rowhammer attack targeting DDR4 servers equipped with Error Correction Code (ECC) memory. Historically, ECC memory has been considered a robust defense against Rowhammer-induced bit flips, either correcting single-bit errors transparently or crashing the system upon detecting multi-bit errors, thereby preventing malicious exploitation. This research, however, demonstrates the first successful Rowhammer attack on DDR4 server DIMMs, effectively bypassing ECC protections to achieve arbitrary bit flips and even forge cryptographic signatures.

Watch on YouTube · Slides

Visual summary for ECC.fail: Mounting Rowhammer Attacks on DDR4 Servers with ECC Memory by Nureddin Kamadan
Visual summary for ECC.fail: Mounting Rowhammer Attacks on DDR4 Servers with ECC Memory by Nureddin Kamadan

Key moments

  1. 0:00 Introduction: Attacking DDR4 servers with ECC memory
  2. 2:00 Understanding the Rowhammer attack mechanism
  3. 2:30 ECC memory as a Rowhammer mitigation on servers
  4. 4:50 Bypassing sampling-based Target Row Refresh (TR) mitigation
  5. 7:00 Recovering the ECC matrix using a logic analyzer
  6. 8:50 Templating victim rows to avoid system crashes

ECC.fail: Mounting Rowhammer Attacks on DDR4 Servers with ECC Memory

Speakers: Nureddin Kamadan

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=MWqtaDhg7Pw

Overview

The "ECC.fail" talk presented at USENIX Security unveils a groundbreaking Rowhammer attack targeting DDR4 servers equipped with Error Correction Code (ECC) memory. Historically, ECC memory has been considered a robust defense against Rowhammer-induced bit flips, either correcting single-bit errors transparently or crashing the system upon detecting multi-bit errors, thereby preventing malicious exploitation. This research, however, demonstrates the first successful Rowhammer attack on DDR4 server DIMMs, effectively bypassing ECC protections to achieve arbitrary bit flips and even forge cryptographic signatures.

The talk, delivered by a presenter identifying himself as Walter, details the complex methodology required to overcome the inherent security of server-grade memory. It highlights critical vulnerabilities in existing Rowhammer mitigations, such as biased Target Row Refresh (TRR) mechanisms, and novel exploitation techniques leveraging obscure CPU features like Chip Queue (CHPQ) error correction capability. The findings presented are highly significant, as they challenge long-held assumptions about the security of server infrastructure and expose a new vector for privilege escalation and data integrity compromise in enterprise environments.

The implications of ECC.fail are profound for data centers, cloud providers, and any organization relying on DDR4 servers. It mandates a re-evaluation of memory security postures and calls for enhanced hardware-level mitigations and sophisticated detection mechanisms. The research not only provides a comprehensive technical deep dive into the vulnerabilities but also outlines a practical, end-to-end attack capable of achieving bit flips in an average of 25 hours and breaking cryptographic signatures in an average of 10 hours, underscoring the real-world feasibility of such exploits.

Background

▶ Watch: Introduction: Attacking DDR4 servers with ECC memory (0:00)

To understand the significance of ECC.fail, it's crucial to grasp the fundamentals of DRAM (Dynamic Random-Access Memory) operation and the Rowhammer vulnerability, as well as the role of ECC in mitigating such threats.

DRAM stores data in a 2D array of capacitors, where the presence or absence of an electrical charge represents a '1' or '0', respectively. The CPU interacts with DRAM by sending specific commands. Two key commands are row activate and refresh. When the CPU needs to read data, it first sends a row activate command, which loads an entire row of capacitors into a row buffer. Data can then be transferred from the buffer to the CPU. Before accessing another row, the currently active row must be pre-charged (or closed), writing its contents back to the DRAM cells. Because capacitors naturally lose charge over time, the CPU must also periodically send refresh commands to the DRAM controller. These commands cause selected rows to be activated and then recharged, ensuring data integrity.

The Rowhammer vulnerability arises from the physical proximity of DRAM cells. Rapidly activating and pre-charging two adjacent rows (known as "aggressor rows") can induce electrical interference in the cells of an unaccessed, "victim" row located between them. This repeated electrical stress can cause the charge in the victim row's cells to leak more quickly than intended. If the charge level drops below the threshold for determining a '1' or '0' before a refresh operation occurs, a bit flip can happen, changing a stored '1' to a '0' or vice-versa. These bit flips, while seemingly random, can be exploited maliciously to achieve privilege escalation (e.g., root access), read/write memory of other processes, or even enable browser-based attacks. Early Rowhammer research primarily focused on consumer-grade computers, where memory lacks advanced error correction capabilities.

Servers, however, typically incorporate ECC (Error Correction Code) memory, designed to enhance data reliability and integrity. ECC memory includes additional bits (check bits) stored alongside the data, allowing the memory controller to detect and, in many cases, correct single-bit errors on the fly. When a single-bit flip occurs, the ECC logic can detect it, correct the erroneous bit, and write the corrected data back to memory. This process is transparent to the operating system and applications. If two or more bits flip within a single ECC codeword (typically a cache line), ECC can detect the multi-bit error but generally cannot correct it. In such critical scenarios, the ECC logic is designed to intentionally crash the machine (a "machine check exception" or "fatal error") to prevent data corruption from propagating or being exploited. This behavior has historically made Rowhammer attacks on ECC-enabled servers seem impractical, as any induced bit flips would either be silently corrected or immediately halt the system, preventing malicious payload execution. The ECC.fail research directly challenges this assumption, demonstrating how to circumvent these robust ECC protections.

Key Findings

▶ Watch: ECC memory as a Rowhammer mitigation on servers (2:30)

The ECC.fail research presents several pivotal findings that collectively enable the first successful Rowhammer attack on DDR4 servers with ECC memory. These discoveries are critical to understanding the feasibility and mechanics of the attack:

  1. First Rowhammer on DDR4 Server DIMMs with ECC: The most significant finding is the successful demonstration of Rowhammer-induced bit flips and subsequent ECC bypass on DDR4 server-grade memory. Previous research largely focused on consumer DIMMs, making this a novel achievement that fundamentally alters perceptions of server memory security. The attack achieves bit flips in an average of 25 hours and can break cryptographic signatures in an average of 10 hours.
  1. Vulnerability in Target Row Refresh (TRR) Mitigation: The researchers discovered that the sampling-based TRR mitigation, employed by DRAM vendors on the tested server DIMMs, is biased. TRR is designed to proactively refresh rows adjacent to frequently accessed aggressor rows to prevent bit flips. However, the observed TRR sampler is significantly more likely to sample row activate commands occurring immediately before a refresh command. This bias allows an attacker to bypass TRR by activating a "dummy" row just before a refresh, causing TRR to refresh rows adjacent to the dummy row, leaving the actual victim row vulnerable to Rowhammer.
  1. Low ECC Code Distance: Through meticulous reverse engineering of the ECC matrix, the researchers found that the ECC scheme used on the tested DDR4 server DIMMs has a surprisingly low code distance of four. This means that if four specific bits within an ECC codeword (a 64-72 byte block) are flipped in a particular pattern, the ECC logic will incorrectly determine that no error has occurred, effectively bypassing the error correction and detection mechanisms. This low code distance is a crucial weakness enabling the bypass.
  1. Timing Side Channel for Single-Bit Flip Detection: To find useful bit flip locations online without crashing the server, the researchers developed a novel detection method. They observed that the Intel memory controller performs DRAM read replays when it encounters data corruption. If a single bit flip occurs, ECC corrects it, but the memory controller first retries the read multiple times to confirm the corruption. This retry mechanism leads to a measurably longer access latency for cache lines containing corrected bit flips compared to clean cache lines. This timing side channel allows the attacker to reliably detect single-bit flips without triggering a machine crash.
  1. Exploitation of Chip Queue (CHPQ) Error Correction Capability: A key enabler for a practical end-to-end attack is the exploitation of a feature called CHPQ (Chip Queue) error correction capability present in Intel CPUs. This capability is designed to keep the server running even if an entire memory chip fails by treating its output as an error that can be corrected by ECC. The researchers discovered that if two specific bits within a 4-bit ECC bypass template are flipped, but these two bits reside on two different physical DRAM chips, the Intel CPU with CHPQ capability misinterprets this as an error originating from a single chip and performs a correction that inadvertently completes the 4-bit bypass template. This effectively reduces the requirement from four difficult-to-achieve flips to just two strategically placed flips, making the attack significantly more practical.

These findings collectively demonstrate that ECC memory, while a strong defense, is not impervious to sophisticated Rowhammer attacks when combined with careful analysis of hardware-specific behaviors and obscure CPU features.

Technical Deep Dive

▶ Watch: Bypassing sampling-based Target Row Refresh (TR) mitigation (4:50)

The ECC.fail attack is a multi-stage process that systematically addresses the challenges of mounting Rowhammer on DDR4 servers with ECC.

The first challenge was getting bit flips on DDR4 server DIMMs. As no prior research existed for DDR4 server DIMMs, the team began by analyzing server DIMMs on an FPGA (Field-Programmable Gate Array). This allowed for precise control over DRAM commands and detailed observation of memory behavior. Their initial observation was the presence of sampling-based Target Row Refresh (TRR), a Rowhammer mitigation. TRR works by randomly sampling one aggressor row from all row activate commands between two refresh commands. When the next refresh command is issued, TRR proactively refreshes the two rows adjacent to the sampled aggressor, aiming to prevent bit flips. However, the critical discovery was that the TRR sampler on these DIMMs was biased. It was significantly more likely to sample row activate commands that occurred immediately before a refresh command. This bias creates a bypass: an attacker can issue a rapid sequence of Rowhammer accesses to a genuine victim row, but then perform a single access to a "dummy" row just before a refresh. The biased TRR logic will likely sample the dummy row, causing it to refresh rows adjacent to the dummy, leaving the actual victim row vulnerable. To confirm the ability to induce flips, the researchers initially disabled ECC checking on their test machine and implemented their Rowhammer program using native code with proper refresh synchronization. This setup successfully yielded more than 100 flips per row.

The second challenge involved recovering the ECC matrix to understand how ECC protects data and, crucially, how to bypass it. ECC works by calculating check bits for data before it's written to memory, which are then stored on an extra chip dedicated to ECC. This process is entirely transparent to software, meaning the check bits cannot be directly read. To overcome this, the researchers used a logic analyzer connected via a DDR4 interposer. This interposer sits between the CPU and memory, allowing the logic analyzer to capture all D-bus traffic, including addresses, commands, data, and the critical check bits. By analyzing the relationship between data and check bits, they reverse-engineered the ECC matrix. Their analysis revealed that the ECC code operates on 64 to 72-byte code sizes and, critically, has a surprisingly low code distance of four. A code distance of four means that if four specific bits within an ECC codeword are flipped in a particular pattern, the ECC logic will mistakenly interpret the data as valid, thus achieving an ECC bypass without triggering an error or correction.

The third challenge was to find useful bit flip locations online without crashing the machine. If the system's ECC was enabled and the researchers hammered with their existing program, it was highly probable that two or more flips would occur within a cache line, immediately crashing the machine. To prevent this, they developed a technique called single-bit templating. They would set the data of the victim row to be identical to the aggressor row, except for a single bit. This setup made that specific bit significantly more likely to flip than others. If this single bit flipped, ECC would correct it, and the machine would not crash. The next hurdle was detecting these corrected single-bit flips, as the data read back would always appear correct. They observed that the Intel memory controller performs DRAM read replays when it encounters data corruption. If the CPU detects a bit flip before ECC corrects it, it will retry the read from that location multiple times to confirm the error. This retry mechanism increases the access latency for cache lines that have experienced a bit flip (even if corrected) compared to those without. This access latency side channel allowed the researchers to infer which specific bits were susceptible to Rowhammer without causing a system crash.

The final and most complex challenge was to build a practical end-to-end attack. Initially, finding four specific bit flips in useful locations to achieve the ECC bypass (given the code distance of four) proved very difficult. The breakthrough came from discovering and exploiting the Chip Queue (CHPQ) error correction capability on Intel server CPUs. This feature is designed to allow the server to continue operating even if an entire DRAM chip provides corrupted data, by correcting the assumed chip-level error. The researchers realized that if they induced only two bit flips, but these two flips occurred in specific locations across two different physical DRAM chips within the 4-bit bypass template, the Intel CPU with CHPQ capability would misinterpret this scenario. It would assume that one of the chips was faulty and proactively "correct" the data by flipping two additional bits in that presumed faulty chip. This unintended correction effectively completed the necessary four-bit pattern for the ECC bypass, requiring only two actual Rowhammer-induced flips from the attacker. This crucial insight significantly reduced the difficulty of achieving a practical ECC bypass.

Demo / Proof of Concept

▶ Watch: Recovering the ECC matrix using a logic analyzer (7:00)

The culmination of the ECC.fail research is a practical, end-to-end Rowhammer attack demonstrated on a live DDR4 server with ECC, leading to the forging of cryptographic signatures. The attack proceeds in several well-defined stages:

  1. Offline Preparations:
  • DMTR and CPU ECC Matrix Recovery: Before the attack on a live system, the researchers perform offline analysis. This involves using an FPGA to understand the DRAM Target Row Refresh (TRR) behavior and identify its biases, as well as employing a logic analyzer with a DDR4 interposer to capture D-bus traffic and reverse-engineer the CPU ECC matrix. This step is crucial for understanding the memory's specific Rowhammer vulnerabilities and ECC layout.
  1. Online Dim Templating (Victim Server):
  • Finding Useful Locations for ECC Bypass: On the target victim server, the attacker uses their Rowhammer program to "template" the DIMM. This involves systematically hammering different rows and employing the single-bit templating technique. By setting victim row data mostly identical to aggressor data, except for one bit, they can induce single-bit flips.
  • Timing Side Channel for Detection: The access latency side channel is used to detect these single-bit flips. A longer access latency indicates a corrected bit flip, revealing locations susceptible to Rowhammer. This process is crucial for identifying the specific bit locations that, when flipped, contribute to the desired 4-bit ECC bypass (or the 2-bit CHPQ-assisted bypass).
  1. Memory Allocator Massaging:
  • Placing Victim Memory: To target specific data (e.g., an RSA public key), the attacker needs to ensure the victim program's critical data is placed on a known vulnerable DRAM row. This is achieved through a technique called DRAM memory allocator massaging.
  • The attacker process first voluntarily frees numerous memory pages that it owns, particularly those corresponding to the identified victim DRAM row(s).
  • Immediately after, the attacker spawns the victim program and temporarily pauses its own execution.
  • The victim program, upon starting, will request memory pages. Operating systems typically maintain a free page pool for each CPU core, containing recently freed pages. The victim program will likely allocate these recently freed pages, which are precisely the pages the attacker just released and which reside on the vulnerable DRAM row.
  • After a short delay, the attacker process resumes execution.
  1. Triggering Rowhammer Flips:
  • With the victim's critical data now residing on the vulnerable DRAM row, the attacker process repeatedly triggers Rowhammer flips on that specific row using the identified aggressor rows and patterns.
  • The attack leverages the CHPQ error correction capability to simplify the bypass: instead of needing to induce four precise flips, only two strategically located flips across different physical chips are required to achieve the effective 4-bit ECC bypass.
  1. Exploitation (RSA Signature Forgery):
  • In the demonstrated proof of concept, the target for the bit flips is an RSA public key stored in the victim's memory.
  • By inducing specific bit flips in the RSA public key's modulus, the attacker can effectively corrupt it in a way that allows for factorization of the flipped public key.
  • Once factored, the attacker can derive the corresponding private key or manipulate signature verification, enabling them to forge RSA signatures.

The researchers report that their end-to-end attack can achieve Rowhammer bit flips in an average of 25 hours and successfully break cryptographic signatures in an average of 10 hours. This demonstrates a practical, albeit time-consuming, method to compromise server integrity via Rowhammer on ECC memory.

Defensive Implications

▶ Watch: Templating victim rows to avoid system crashes (8:50)

The ECC.fail research exposes critical vulnerabilities in the memory security of DDR4 servers, necessitating a multi-faceted defensive strategy. Defenders, including DRAM vendors, CPU manufacturers, system integrators, and cloud providers, must consider the following implications:

  1. Improve Target Row Refresh (TRR) Mitigations: The discovery of a biased sampling-based TRR mechanism is a significant finding. DRAM vendors must re-evaluate and enhance their TRR implementations to ensure unbiased sampling and more robust protection against Rowhammer. This might involve firmware updates for DIMMs or new hardware designs. The TRR logic should not be bypassable by specific access patterns, especially those involving "dummy" rows.
  1. Increase ECC Code Distance and Robustness: The observed low ECC code distance of four is a fundamental weakness. While increasing code distance adds overhead, it significantly improves resilience against multi-bit errors. Future ECC designs should aim for higher code distances or implement more sophisticated error detection and correction algorithms that are less susceptible to structured multi-bit flips. Furthermore, the reliance on single-bit correction to mask initial signs of Rowhammer should be re-evaluated.
  1. Address CHPQ Vulnerabilities: The exploitation of Chip Queue (CHPQ) error correction capability highlights how seemingly benign features designed for reliability can be repurposed for attack. CPU manufacturers, particularly Intel, need to investigate whether microcode updates or architectural changes can mitigate this specific misinterpretation of multi-chip errors that leads to an inadvertent ECC bypass. The logic for multi-chip error handling needs to be scrutinized for potential attack vectors.
  1. Enhanced Memory Allocator Security: The memory allocator massaging technique demonstrates how an attacker can manipulate memory placement to target specific data. Operating systems and hypervisors could implement more robust memory randomization or allocate sensitive data with higher isolation guarantees to prevent an attacker from reliably predicting and influencing the physical location of victim data on DRAM rows.
  1. Monitor Memory Access Latency: The access latency side channel used to detect single-bit flips could also be a defensive mechanism. Security monitoring solutions could potentially track unusual patterns of increased memory access latencies, which might indicate active Rowhammer templating attempts, even if ECC is correcting the underlying bit flips. However, this would require high-resolution, low-level monitoring capabilities not typically available to standard enterprise monitoring tools.
  1. Hardware-Level Protections and Isolation: Longer-term solutions may involve architectural changes to DRAM, such as physical isolation between rows, or more sophisticated on-die monitoring and mitigation techniques. For cloud environments, stricter physical memory isolation between virtual machines is paramount to prevent one tenant from hammering another.
  1. Regular Firmware and Microcode Updates: System administrators should prioritize applying firmware updates for DIMMs and microcode updates for CPUs as they become available, as these updates may contain patches for Rowhammer-related vulnerabilities or improvements to ECC and TRR mechanisms.

In summary, the ECC.fail attack necessitates a collaborative effort across the hardware and software stack to re-establish the integrity of server memory. Relying solely on current ECC implementations is no longer sufficient to guarantee protection against advanced Rowhammer attacks.

Key Takeaways

  • ECC is Not Impenetrable: DDR4 servers with ECC memory are not immune to Rowhammer attacks. This research demonstrates the first successful Rowhammer attack capable of bypassing ECC protections on server-grade hardware.
  • Biased TRR is a Critical Flaw: The Target Row Refresh (TRR) mitigation, designed to prevent Rowhammer, was found to have a significant bias that allows attackers to bypass it by strategically accessing "dummy" rows.
  • Low ECC Code Distance is an Enabler: The surprisingly low code distance of four in the tested ECC implementation means that specific patterns of four bit flips can completely bypass error detection and correction.
  • CHPQ Feature Exploitation: A key to practical exploitation is the abuse of Intel's Chip Queue (CHPQ) error correction capability, which allows two strategically placed bit flips across different chips to be misinterpreted and "corrected" into a full 4-bit ECC bypass.
  • Timing Side Channels Reveal Flips: Even when ECC corrects single-bit flips, the increased access latency due to DRAM read replays can be used as a side channel to identify vulnerable memory locations without crashing the system.
  • Practical End-to-End Attack: The research demonstrates a complete attack chain, including memory allocator manipulation and RSA signature forgery, achieving bit flips in 25 hours and signature breaks in 10 hours on average.

About the Speaker(s)

The talk "ECC.fail: Mounting Rowhammer Attacks on DDR4 Servers with ECC Memory" was presented by Walter, who introduced himself at the beginning and conclusion of the presentation. The metadata for the talk lists Nureddin Kamadan as the speaker. While the transcript provides no specific biographical details about Walter beyond his name, the work presented is a detailed technical exposition on memory security, indicating expertise in hardware security, DRAM internals, and exploitation techniques. The mention of Nureddin Kamadan in the metadata suggests he is a key contributor or lead author of the research.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This is the real thing: original, difficult hardware security research that breaks a long-standing assumption the industry has been leaning on for years. Bypassing ECC on DDR4 server DIMMs via a combination of biased TRR exploitation, ECC matrix reverse engineering, a timing side-channel for stealthy flip detection, and abuse of Intel's CHPQ feature is not a talk — it's a multi-year research program compressed into a single session. The end-to-end RSA signature forgery demo seals it.

Heather Calloway (CISO) — WEAK

Rigorous, technically impressive research that cracks a foundational assumption about server memory security — but the presentation never crosses the line from 'interesting to researchers' to 'actionable for anyone who runs infrastructure.' The defensive section lists hardware vendor wish list items that operators cannot act on, and the institutional exposure question — who is accountable, which cloud providers are affected, what is the risk posture for enterprises today — goes unaddressed.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)