"I'm regretting that I hit run": In-situ Assessment of Potential Malware

Brandon Lit

34th USENIX Security Symposium (USENIX Security '25) · Day 2 · Software Security and Usable Security

Overview

This article delves into "FLOP: Breaking the Apple M3 CPU via False Load Output Predictions," a significant research paper presented at USENIX Security. The work, authored by Jason Kim, Jalen Chuang, and Daniel Genkin from Georgia Tech, and Yuval Yarom from Ruhr University Bochum, uncovers a novel microarchitectural vulnerability stemming from a Load Value Predictor (LVP) implemented in recent Apple M- and A-series processors, including the M3, M4, and A17 Pro. This LVP, a performance optimization designed to mitigate Read-After-Write (RAW) dependencies, is shown to introduce critical security flaws by speculatively computing on stale or incorrect data values.

Read the paper · Download the PDF (PDF) · Slides

Paper abstract

To bridge the ever-increasing gap between the fast execution speed of modern processors and the long latency of memory accesses, CPU vendors continue to introduce newer and more advanced optimizations. While these optimizations improve performance, research has repeatedly demonstrated that they may also have an adverse impact on security. In this work, we identify that recent Apple M- and A-series processors implement a load value predictor (LVP), an optimization that predicts the contents of memory that the processor loads before the contents are actually available. This allows processors to alleviate slowdowns from Read-After-Write dependencies, as instructions can now be executed in parallel rather than sequentially. To evaluate the security impact of Apple's LVP implementation, we first investigate the implementation, identifying the conditions for prediction. We then show that although the LVP cannot directly predict 64-bit values (e.g., pointers), prediction of smaller-size values can be leveraged to achieve arbitrary memory access. Finally, we demonstrate end-to-end attack exploit chains that build on the LVP to obtain a 64-bit read primitive within the Safari and Chrome browsers.

Visual summary for "I'm regretting that I hit run": In-situ Assessment of Potential Malware by Brandon Lit
Visual summary for "I'm regretting that I hit run": In-situ Assessment of Potential Malware by Brandon Lit

FLOP: Breaking the Apple M3 CPU via False Load Output Predictions

Speakers: Jason Kim (Georgia Tech); Jalen Chuang (Georgia Tech); Daniel Genkin (Georgia Tech); Yuval Yarom (Ruhr University Bochum)

Conference: USENIX Security

YouTube: https://www.usenix.org/conference/usenixsecurity25/presentation/kim-jason

Overview

This article delves into "FLOP: Breaking the Apple M3 CPU via False Load Output Predictions," a significant research paper presented at USENIX Security. The work, authored by Jason Kim, Jalen Chuang, and Daniel Genkin from Georgia Tech, and Yuval Yarom from Ruhr University Bochum, uncovers a novel microarchitectural vulnerability stemming from a Load Value Predictor (LVP) implemented in recent Apple M- and A-series processors, including the M3, M4, and A17 Pro. This LVP, a performance optimization designed to mitigate Read-After-Write (RAW) dependencies, is shown to introduce critical security flaws by speculatively computing on stale or incorrect data values.

The researchers meticulously characterize Apple's LVP, demonstrating that while it doesn't directly predict 64-bit pointer values, it can be leveraged through indirection to achieve arbitrary memory access. They present two end-to-end attack exploit chains: FLOP-Data targeting Apple Safari and FLOP-Control targeting Google Chrome. Both attacks culminate in obtaining a 64-bit read primitive, allowing for sandbox escapes and the exfiltration of sensitive user data from popular web services, such as location history from Google Maps, email content from Proton Mail, and credit card information from Square storefronts. This research highlights a new class of microarchitectural attack specific to Apple's modern silicon, underscoring the ongoing challenge of balancing performance optimizations with robust security in contemporary CPU designs.

Background

Modern CPUs constantly strive to bridge the performance gap between their fast execution units and the comparatively slow memory subsystems. To achieve this, processors employ a myriad of optimizations, including caches, out-of-order execution, and speculative execution. While these techniques significantly boost performance, they have also been repeatedly shown to introduce security vulnerabilities. Landmark attacks like Meltdown [31] and Spectre [27] demonstrated how speculative execution, where the CPU predicts future instruction paths, can leave exploitable microarchitectural side effects even if the speculative execution is later rolled back. These attacks have since spawned numerous follow-ups, challenging nearly all hardware-backed security domains.

A fundamental challenge in CPU pipelining is dealing with Read-After-Write (RAW) dependencies. These occur when a younger instruction needs to read from a memory location that an older, still-executing instruction is writing to. Traditionally, the CPU must stall the younger instruction until the older write completes, serializing execution and introducing performance bottlenecks. To alleviate this, computer architects have proposed value prediction, where the CPU attempts to predict the outcome of an instruction before it's actually computed, allowing dependent instructions to proceed speculatively. A common form of value prediction is the Load Value Predictor (LVP), which specifically targets load operations. LVPs observe values returned from loads and, if consistent patterns are detected (e.g., a load repeatedly returns the same value), they can predict that value for subsequent loads. This allows the CPU to speculatively execute instructions that depend on the predicted load value in parallel with the actual memory access. If the prediction is correct, performance is enhanced; if incorrect, the speculative work is rolled back, and execution is replayed. However, as this research demonstrates, the transient computation performed during a misprediction can leave detectable traces, forming the basis for powerful side-channel attacks. The exfiltration of these traces often relies on well-known cache side-channel attacks like FLUSH+RELOAD [17, 62] and PRIME+PROBE [13, 22, 32, 38, 43, 44, 58], which monitor cache state changes to infer secret information.

Key Findings

The research makes several critical discoveries regarding the Apple LVP and its security implications:

  • LVP Discovery and Presence: A novel Load Value Predictor (LVP) was identified and characterized on recent Apple M-series (M3, M4) and A-series (A17 Pro) CPUs. Notably, this LVP is absent on earlier generations like the M2, A15 Bionic, and A16 Bionic processors.
  • Activation Criteria: The Apple LVP activates only for constant load values, demonstrating no prediction capability for striding or randomly changing values. It maintains training state on a per-instruction address basis, with an observed capacity to track up to 72 different instruction addresses simultaneously.
  • Load Width Limitations: The LVP effectively predicts arbitrary values for 4-byte-wide loads and smaller. Crucially, for 8-byte loads, which typically represent 64-bit pointers on these architectures, prediction only occurs if the load value is zero. This appears to be a deliberate countermeasure against direct pointer prediction, though the researchers show it can be bypassed.
  • Training and Speculation Window: Approximately 250 training loads are sufficient to reliably train the LVP and induce mispredictions. Once triggered, an LVP misprediction can lead to a speculation window lasting up to 330 cycles when the target memory is not cached (a cache miss), and around 30 cycles for cache hits.
  • Microarchitectural Side Effects: LVP mispredictions cause transient computation on stale load values. This can be exploited to violate memory safety, leading to 64-bit out-of-bounds reads and diversion of control flow to rogue functions.
  • End-to-End Exploits: Two practical attack chains were developed:
  • FLOP-Data (Safari): Achieves speculative type confusion in WebKit, leading to 64-bit out-of-bounds reads and exfiltration of sensitive user data from cross-origin webpages (e.g., Google Maps location history, Proton Mail inbox, iCloud Calendar events).
  • FLOP-Control (Chrome): Exploits a double LVP misprediction to transiently execute the wrong WebAssembly function with unchecked arguments, effectively treating a 64-bit integer as a memory address to read sensitive data (e.g., credit card information and billing addresses from Square storefronts).
  • Kernel-Space Implications: The LVP functions similarly in the macOS kernel, demonstrating the potential for in-kernel LVP training and exploitation, although such attacks require specific kernel gadgets.
  • Mitigation through DIT: The Data Independent Timing (DIT) bit, a feature in Armv8.4-A and newer ISAs, was found to effectively disable the LVP on the M3 CPU on a per-process basis, requiring no special privileges to set. Enabling DIT introduces a minimal performance overhead (e.g., 4.5% on Speedometer 3.0 for Safari, 0.6% for native benchmarks).

Technical Deep Dive

The core of this research involved a meticulous reverse engineering and characterization of the Apple LVP. The initial discovery was made using a carefully crafted experiment (Listing 1) that measured memory access latency on randomly shuffled addresses. The key insight was to introduce a Read-After-Write (RAW) dependency by accumulating load values into a junk variable, forcing serialization. By varying the mem buffer content between constant values (experiment) and random values (control), the researchers observed a significant speedup on the M3's P-cores (Figure 3, top right) when load values were constant, indicating the presence of an LVP. This speedup was not observed on the M2 CPU or the E-cores of either M2 or M3, confirming the LVP's specific presence on newer performance cores.

Further investigation into LVP behavior revealed crucial details:

  • LVP Presence on Other Apple CPUs: A portable WebAssembly version of the experiment confirmed LVP presence on M3, M4, and A17 Pro, and its absence on M2, A15 Bionic, and A16 Bionic (Table 1).
  • Load Width and Value Prediction: The LVP predicts arbitrary values for 1, 2, and 4-byte loads. However, for 8-byte loads (64-bit pointers), prediction only occurs if the load value is zero (Figure 4). This suggests a hardware-level countermeasure to prevent direct leakage of pointer values, though the researchers bypass this with indirection.
  • Prediction of Constant Values: The LVP trains exclusively on constant load values. When presented with striding values, the runtime resembled that of random values, indicating no LVP activation (Figure 5).
  • Instruction Address Tagging: Loop unrolling, which assigns unique instruction addresses to each load, prevented LVP activation (Figure 6). This strongly suggests that the LVP's training state is tagged with the instruction address (PC-tagging) rather than operating on a global window.
  • Simultaneous Prediction Capacity: The LVP can track and activate on up to 72 distinct load instructions, although this capacity varies with the instruction distance between them (Figure 7). The researchers conjecture a 4-way set-associative internal state cache, hashing page offset bits.

To demonstrate LVP-induced speculative computation, the researchers designed a gadget (Listing 2) that trains the LVP on a value (foo), then changes the architectural value to another (bar), and observes the transient foo through a FLUSH+RELOAD covert channel (Figure 8).

  • Activation Threshold: Reliable mispredictions were achieved after approximately 250 training loads. The researchers observed spikes in activation rates around 60, 120, and 180 loads before reaching near-perfect reliability past 240 loads (Figure 9).
  • State Persistence: The LVP's internal state is remarkably persistent, enduring over one-second busy waits (more than four billion cycles) without concurrent memory-intensive workloads. However, intensive load/store activity can overwrite the state over time, and parking the CPU core (e.g., via sleep) resets it (Figure 10).
  • Speculation Depth: The speculation window, measured by inserting mul instructions between the speculative load and covert channel transmission, extends up to 330 cycles for cache misses (110 mul instructions, each taking 3 cycles) and 30 cycles for cache hits (10 mul instructions) (Figure 11).

Crucially, the researchers showed how to leverage these mispredictions to achieve memory safety violations despite the LVP's 8-byte load limitation. By modifying the gadget (Listing 3, Figure 12), the stale LVP value (foo) was used as an index into an array of pointers (aop), allowing an attacker to select and dereference an incorrect pointer. This enabled 64-bit out-of-bounds reads, achieving a mean accuracy of 0.97 (median 1.00) and a throughput of 210,526 bits per second. Furthermore, by making aop an array of function pointers, the researchers demonstrated control flow diversion to rogue functions, achieving a rate of approximately 2,551 rogue function calls per second (Figure 13). These experiments were also successfully replicated in the macOS kernel, demonstrating similar accuracy and throughput, highlighting the severity of the vulnerability in a privileged context.

Demo / Proof of Concept

The researchers developed two end-to-end attacks, FLOP-Data and FLOP-Control, demonstrating the practical exploitability of LVP mispredictions in real-world browser environments.

FLOP-Data: Speculative Type Confusion in Safari

The FLOP-Data attack targets Apple Safari by exploiting speculative type confusion in WebKit, its underlying browser engine. JavaScript, being weakly typed, relies on WebKit to perform runtime type checks. Every JavaScript data structure in WebKit starts with a JSCell header (Figure 14), where the typeVar (4 bytes) is loaded first (Listing 4, Line 1). Since this is a 4-byte load, the LVP can be trained to predict the EXPECTED_TYPE.

The attack leverages a scenario where the typeVar and the backingStore (or optionalField) are on different cache lines. While Apple patched Intl.Locale objects to prevent this split, the researchers found that typed arrays (e.g., Uint8Array) can be allocated such that their Type, Misc., and BS (backing store) are on one cache line, but rawBuf (the address to their raw buffer) and other metadata are on another (Figure 15). The attack then aims to confuse a malicious typed array (evil) with a small JavaScript object containing a string (objWithStr). Small objects store their properties inline, and prop (the string object) overlaps with rawBuf of the typed array (Figure 16).

To achieve 64-bit reads, the rawBuf of evil is crafted to resemble the header of a legitimate string object. This involves writing the constant string type (0x4250) and miscellaneous data (0x80200) into evil's rawBuf (Figure 17). The StrData field of the forged string, which normally points to its underlying data, is set to an attacker-controlled address (e.g., 0x400000100). To bypass WebKit's pointer poisoning and provide a valid target for StrData, the researchers employ a heap spray technique. They create a 1 GiB JavaScript string containing multiple copies of a forged string data structure (Figure 18). This forged data structure has a large length (to bypass bounds checks) and an attacker-controlled charPtr (the target addr for the 64-bit read). This large string allocation consistently lands at predictable addresses around 0x400000000, making 0x400000100 a reliable target.

The gadget (Listing 5) input.prop.charCodeAt(index) is used to trigger the attack. The LVP is trained with objWithStr, causing it to predict the correct string type. When evil is then passed as input, its typeVar is flushed. The LVP mispredicts the type, causing the CPU to speculatively interpret evil's rawBuf as objWithStr's prop. The forged string header in rawBuf passes the type check, and the CPU dereferences the fakeStrData (0x400000100), leading it to the sprayed forged string data. Finally, the charCodeAt function, operating on the forged string, uses the attacker-controlled charPtr (addr) and index to perform a 64-bit dereference, leaking the data via a microarchitectural covert channel (Figure 19).

Benchmarking FLOP-Data on an M3 MacBook Pro yielded a median accuracy of 89.58% and a throughput of 0.492 bits per second. For end-to-end impact, the window.open API was used to co-render target webpages (e.g., Google Maps, Proton Mail, iCloud Calendar) in the same Safari process, allowing the attack to access their DOM data. The researchers demonstrated leakage of sensitive information, including location history, email sender/subject, and private calendar events (Figure 20).

FLOP-Control: Control Flow Hijacking in Google Chrome

The FLOP-Control attack targets Google Chrome, specifically exploiting WebAssembly's function dispatch table (Figure 21) to achieve speculative control flow hijacking and 64-bit reads. Chrome cages JavaScript objects in a 4 GiB region with 32-bit pointers, but WebAssembly code is not subject to this restriction. WebAssembly functions can be inserted into a dispatch table and invoked by index. Each table entry (funcData) contains an entryPt (code pointer) and sig (32-bit function signature).

The attack leverages the call function (Listing 6), which first loads funcData using an index (Line 2), then loads sig (Line 3), and performs a funcSig check (Lines 4-5) before calling entryPt (Line 7). Both index and sig are 32-bit values, making them predictable by the LVP. The challenge is that a misprediction of index (leading to the wrong funcData) would cause the funcSig check to fail, aborting speculation.

To bypass this, the researchers orchestrate a double LVP misprediction (Figure 22). They find a funcData entry where entryPt and sig are on different cache lines, and index and sig are not cached.

  1. First Misprediction: When call(args, 0) is invoked, the LVP mispredicts index to be 1 (a stale value), causing the CPU to speculatively retrieve the funcData of wrongFunc instead of func.
  2. Second Misprediction: Subsequently, the load for sig also misses the cache. The LVP again mispredicts, this time predicting the sig to be EXPECTED_SIG (the signature of func). This bypasses the funcSig check.
  3. Control Flow Diversion: With the checks bypassed, the CPU speculatively sets up func's arguments for wrongFunc. Since wrongFunc's entryPt is cached, control flow is diverted to wrongFunc.

The key to 64-bit reads lies in confusing data as addresses. While WebAssembly doesn't expose raw pointers, Chrome implements WebAssembly struct references as 64-bit pointers. The attack defines func(uint64_t arg) and wrongFunc(struct readType* arg) (Listing 7). By calling func with a 64-bit integer, but having the LVP misdirect to wrongFunc, the CPU speculatively treats the integer argument as a 64-bit pointer to readType, dereferencing it to leak data via a covert channel.

Benchmarking FLOP-Control on an M3 MacBook Pro resulted in a median accuracy of 80.90% and a throughput of 0.30 bits per second. To target real-world secrets, the attack leverages Chrome's eTLD+1 site isolation rule. While this usually prevents co-rendering of different origins, the researchers found that Square storefronts (e.g., attacker.square.site) are not on the Public Suffix List, allowing them to co-render with other Square domains (e.g., customer-account.square.site) using window.open. This enables access to the customer account page's DOM, which contains saved credit card information and billing addresses (Figure 23, Left). After defeating ASLR (taking ~30 seconds and reducing address entropy to 12 bits), FLOP-Control successfully exfiltrated the last 4 credit card digits, expiration date, and billing address (Figure 23, Right).

Defensive Implications

The discovery of the LVP vulnerability necessitates immediate defensive measures to protect user data and maintain system integrity. The researchers identify a promising mitigation: the Data Independent Timing (DIT) bit. Present in Armv8.4-A ISA and newer, the DIT bit instructs the CPU to ensure that instruction latency does not correlate with operand data, a feature originally designed to protect constant-time code. The paper confirms that setting the DIT bit on the M3 CPU effectively disables the LVP, preventing both speedups from constant loads and mispredictions.

Recommendations for Developers and Vendors:

  • Enable DIT Bit: Developers should patch software to enable the DIT bit on supported platforms, particularly for code regions handling secrets or executing untrusted code (e.g., user-supplied JavaScript or WebAssembly in browsers, password fields in applications).
  • The overhead of enabling DIT is minimal: a patched Safari browser showed a 4.5% overhead on the Speedometer 3.0 benchmark, and native environments experienced an average of 0.6% overhead on the BYTE Unix benchmark. This low performance cost makes DIT a highly viable mitigation strategy.
  • Browser-Specific Mitigations:
  • Safari: The cross-origin leaks in Safari were facilitated by WebKit's lack of Site Isolation, which allows attacker and target webpages to be rendered in the same address space. Implementing robust Site Isolation (similar to Chrome's model) is crucial. Additionally, increasing the randomization entropy for JavaScript object types and memory allocations would hinder attackers from forging data structures (like the fake string header) and reliably performing heap sprays.
  • Chrome: While Chrome's Site Isolation is more robust, corner cases like the Square attack demonstrate its limitations. To address the WebAssembly struct reference vulnerability, Chrome should cage WebAssembly structs within a memory region, similar to its existing protection for JavaScript objects. This would restrict references to base-relative 32-bit offsets instead of 64-bit pointers, preventing arbitrary address dereferences.
  • Kernel Hardening: The LVP is enabled in the macOS kernel, posing a significant risk if suitable gadgets are found. While the researchers couldn't train the LVP across userspace/kernel boundaries, future work is needed to find kernel-space training, converter, and leakage gadgets to fully assess and mitigate this risk.

Key Takeaways

  • Novel Microarchitectural Vulnerability: Recent Apple M- and A-series CPUs (M3, M4, A17 Pro) implement a Load Value Predictor (LVP) that, while enhancing performance, introduces a new class of microarchitectural security vulnerability.
  • Exploitable Speculative Execution: The LVP can be reliably trained to mispredict load values, leading to transient computation on stale data within a speculation window of up to 330 cycles.
  • Arbitrary Memory Access: Despite LVP not directly predicting 64-bit pointers, it can be leveraged through indirection (e.g., using stale values as array indices) to achieve 64-bit out-of-bounds reads and control flow hijacking.
  • Real-World Browser Exploits: End-to-end attacks, FLOP-Data for Safari and FLOP-Control for Chrome, demonstrate practical sandbox escapes and cross-origin data exfiltration of sensitive information like location history, emails, and credit card details.
  • Effective Mitigation Available: The Arm Data Independent Timing (DIT) bit can effectively disable the LVP with minimal performance overhead (e.g., 4.5% in Safari), offering a practical defense for sensitive code execution.
  • Continuous Security Challenge: This research underscores the ongoing challenge for CPU designers and software developers to anticipate and mitigate security implications of performance optimizations, especially in the evolving landscape of ARM-based CPUs.

About the Speaker(s)

The research presented in "FLOP: Breaking the Apple M3 CPU via False Load Output Predictions" was a collaborative effort by security researchers from leading academic institutions. Jason Kim and Jalen Chuang are affiliated with Georgia Tech, where they contributed as joint first authors to this significant work. They are joined by Daniel Genkin, also from Georgia Tech, and Yuval Yarom from Ruhr University Bochum. Their collective expertise spans hardware security and microarchitectural vulnerabilities, contributing to a deeper understanding of the security landscape of modern processors.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This is the real deal — novel hardware vulnerability research on Apple's newest silicon, reverse-engineered from scratch, with two end-to-end browser exploits leaking actual user data. The kind of work that forces Apple to patch and makes every other CPU vendor sweat about their own value predictors.

Heather Calloway (CISO) — STRONG ACCEPT

High-quality microarchitectural research demonstrating practical cross-origin data exfiltration on Apple M3/M4 silicon via browser-based attacks. Any CISO with an Apple fleet—particularly executive protection, BYOD, or high-value target populations—needs to understand this exists and track mitigation deployment.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)