Repurposing Neural Networks for Efficient Cryptographic Computation
Xin Jin
Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Privacy & Cryptography 1 · Privacy & Cryptography 1
Overview
In an era increasingly reliant on robust digital security, the performance of cryptographic operations remains a critical bottleneck. This talk, "Repurposing Neural Networks for Efficient Cryptographic Computation," presented by Xin Jin at the NDSS Symposium, introduces TensorCrypt, a novel framework that leverages the inherent computational graph properties of neural networks to significantly accelerate cryptographic computations. Instead of traditional machine learning approaches that attempt to train a model for cryptographic functions, TensorCrypt directly transforms cryptographic algorithms into neural network computational graphs, effectively repurposing existing AI hardware and software stacks for high-speed encryption and decryption.
Key moments
- 0:00 Introduction and problem: accelerating crypto computation
- 2:00 Limitations of existing solutions; neural networks' promise
- 4:00 Why direct neural network training for crypto fails
- 5:00 TensorCrypt's core insight: leveraging computational graphs
- 6:00 TensorCrypt framework, transformation rules, and formal proof
- 7:15 Key optimizations: operator fusion and load-on-use management
- 9:10 Evaluation results: TensorCrypt is up to 5.4x faster
Repurposing Neural Networks for Efficient Cryptographic Computation
Speakers: Xin Jin
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=mtsUu0bOyjM
Overview
In an era increasingly reliant on robust digital security, the performance of cryptographic operations remains a critical bottleneck. This talk, "Repurposing Neural Networks for Efficient Cryptographic Computation," presented by Xin Jin at the NDSS Symposium, introduces TensorCrypt, a novel framework that leverages the inherent computational graph properties of neural networks to significantly accelerate cryptographic computations. Instead of traditional machine learning approaches that attempt to train a model for cryptographic functions, TensorCrypt directly transforms cryptographic algorithms into neural network computational graphs, effectively repurposing existing AI hardware and software stacks for high-speed encryption and decryption.
The motivation behind TensorCrypt stems from the pervasive need for efficient cryptography in modern applications, such as end-to-end encrypted video conferencing (e.g., Zoom), which demands processing vast amounts of data—upwards of 5 megabytes per second for 50 participants, translating to 260,000 AES blocks per second. While dedicated hardware like AS-NI or general-purpose accelerators like GPUs and FPGAs exist, they often suffer from hardware diversity, complex software development, and suboptimal software optimization. TensorCrypt addresses these challenges by tapping into the highly optimized, widely supported, and hardware-agnostic ecosystem of neural network frameworks and accelerators.
This research is significant because it offers a paradigm shift in how cryptographic performance can be achieved. By treating cryptographic algorithms as computational graphs that can be mapped directly onto neural network structures, TensorCrypt unlocks the potential of specialized AI hardware (like Nvidia GPUs and Google TPUs) and optimized deep learning compilers (like TVM and XLA) for security-critical tasks. The framework promises not only substantial speedups but also enhanced deployment flexibility across a wide range of devices, from cloud servers to IoT devices, without compromising the rigorous security guarantees of the underlying cryptographic primitives.
Background
▶ Watch: Introduction and problem: accelerating crypto computation (0:00)
The demand for efficient cryptographic computation is escalating rapidly with the proliferation of secure communication and data storage. Applications like end-to-end encrypted video calls, secure cloud storage, and privacy-preserving computations all rely on cryptographic primitives that, while mathematically robust, can be computationally intensive. For instance, securing a Zoom call with 50 participants requires encrypting and decrypting approximately 5 megabytes of data per second. If AES is used, this translates to processing around 260,000 blocks per second – a formidable computational load.
Historically, efforts to accelerate cryptographic operations have fallen into several categories. Instruction Set Extensions like AES-NI (Advanced Encryption Standard New Instructions) on Intel and AMD CPUs provide dedicated hardware instructions for specific ciphers, offering significant speedups but tying implementations to specific CPU architectures. General-Purpose Graphics Processing Units (GPUs) and Field-Programmable Gate Arrays (FPGAs) have also been widely adopted for cryptographic acceleration. GPUs, with their massive parallel processing capabilities, are excellent for highly parallelizable tasks, including certain cryptographic modes. FPGAs offer even greater flexibility and efficiency through custom hardware logic.
However, these existing solutions present several limitations. They often require diverse and specialized hardware and software stacks, leading to complex development cycles and maintenance overhead. Furthermore, many GPU/FPGA-based accelerations primarily leverage hardware capabilities without extensive software optimization, potentially leading to suboptimal performance. Developers must contend with low-level programming models (e.g., CUDA for GPUs) and hardware-specific configurations, complicating deployment across heterogeneous environments.
In parallel, the rise of large language models (LLMs) and the broader field of deep learning have propelled neural networks into the forefront of computational efficiency. Neural networks are supported by a vast ecosystem of hardware accelerators (e.g., Nvidia GPUs, Google TPUs), diverse software frameworks (e.g., PyTorch, TensorFlow), and highly optimized compilers (e.g., TVM, XLA). They are designed for efficient tensor manipulation and can be deployed across a wide spectrum of devices, from high-performance servers to resource-constrained IoT devices. Crucially, neural networks are Turing complete, meaning they can theoretically compute any computable function.
An intuitive, yet ultimately flawed, approach to accelerating cryptography with neural networks would be to train a model to learn cryptographic functions. However, this is highly challenging. Cryptographic algorithms often involve high nonlinearity, large (or infinite) input/output spaces, and require exact, deterministic outputs—a single bit error renders the entire operation invalid. Traditional neural networks, especially those trained on data, are probabilistic and approximate; they struggle with such strict requirements and the immense data needed to cover the vast input-output space of a cryptographic function. State-of-the-art models like OpenAI's GPT-3 would fail to produce correct ciphertexts if simply trained on cryptographic data due to these fundamental mismatches. This observation led the TensorCrypt researchers to avoid a training-based approach entirely.
Key Findings
▶ Watch: Why direct neural network training for crypto fails (4:00)
The core finding of the TensorCrypt research is that cryptographic computations can be directly represented and executed as neural network computational graphs without any training. This insight allows the repurposing of existing, highly optimized neural network hardware and software stacks for efficient cryptographic acceleration. The key contributions include:
- A Novel Program Transformation Framework: TensorCrypt introduces a formal framework that transforms cryptographic algorithms, expressed in a Domain Specific Language (DSL), into equivalent neural network computational graphs. This transformation ensures semantic and security equivalence, meaning the transformed neural network performs the exact cryptographic operation and does not leak additional information.
- Model Optimizations for Performance: The framework incorporates specific optimizations tailored for cryptographic workloads. These include operator fusions to reduce overheads from data transformation and operator scheduling, and a novel load-on-use memory management strategy to minimize memory swapping on accelerators like GPUs, particularly for frequently accessed but intermittently used components like counters in block cipher modes.
- Significant Performance Gains: Through extensive evaluation, TensorCrypt demonstrates substantial speedups, achieving up to 5.4 times faster cryptographic computation compared to highly optimized GPU-based baselines. This performance is validated across critical ciphers like AES, Chacha, and Salsa, which are widely used, with AES and Chacha being the only two ciphers selected in TLS 1.3 for large-volume data encryption/decryption.
- Hardware and Software Stack Agnostic Deployment: The transformed neural network models can be deployed across a diverse range of hardware platforms (Nvidia GPUs, Google TPUs, CPUs) and software environments (cloud, desktop, laptop, smartphone, IoT devices), leveraging the inherent flexibility and optimization of the neural network ecosystem. This broad compatibility addresses the limitations of hardware-specific acceleration solutions.
Technical Deep Dive
▶ Watch: TensorCrypt's core insight: leveraging computational graphs (5:00)
The fundamental premise of TensorCrypt lies in a profound observation: both neural networks and cryptographic algorithms can be conceptualized as computational graphs. A computational graph represents a series of operations applied to data, where nodes are operations (or operators) and edges are data dependencies (tensors). Neural networks are inherently computational graphs, processing data in the form of tensors through layers of mathematical operations. Similarly, cryptographic algorithms, at their core, consist of well-defined sequences of mathematical and logical operations on input data (plaintext, keys, nonces) to produce output data (ciphertext).
TensorCrypt proposes a program transformation framework to bridge these two domains. The framework begins by abstracting both cryptographic computations and neural network structures into Domain Specific Languages (DSLs). This abstraction provides a high-level, formal representation that facilitates systematic analysis and transformation.
The transformation process itself operates at two distinct layers:
- Program Layer Transformation: At this higher level, the entire cryptographic program is analyzed and transformed into a holistic neural network computational graph. This involves identifying the overall data flow and control structures of the cryptographic algorithm and mapping them to an equivalent, single neural network model.
- Statement Layer Transformation: Below the program layer, individual statements or elementary operations within the cryptographic algorithm are transformed into corresponding neural network layers or subgraphs. For example, a bitwise XOR operation, a modular addition, or a substitution box (S-box) lookup can be represented using specific tensor operations (e.g., element-wise XOR, matrix multiplication, lookup tables implemented as tensors).
Crucially, the framework includes semantic rules that govern these transformations. While not detailed in the talk due to time constraints, these rules are essential for ensuring the correctness and security of the transformed output. They formally define how cryptographic operations map to tensor operations, preserving the exact mathematical logic of the original algorithm.
A cornerstone of TensorCrypt is its formal security proof. The researchers rigorously prove that the transformation process does not introduce any new security vulnerabilities or leak additional information to attackers. The transformed neural network model is semantically and security-equivalent to the original cryptographic implementation. This means that if the original cryptographic algorithm is secure, its TensorCrypt-transformed counterpart maintains the same security level. This addresses critical concerns about potential cryptographic weaknesses introduced by the transformation, such as those that might arise from minor protocol changes. The speaker explicitly noted that the design is at the same security level, though the implementation is different. While the design is equivalent, the speaker acknowledged that side-channel attacks (e.g., timing attacks) against the implementation are still a possibility, much like with any software implementation. However, the speaker clarified that adversarial machine learning attacks (e.g., perturbing inputs to cause misclassification) are not applicable to TensorCrypt models because they are not learning-based; they are direct, deterministic constructions of cryptographic ciphers using computational graphs.
Beyond the core transformation, TensorCrypt incorporates several model optimizations to maximize performance on neural network accelerators:
- Operator Fusions: This optimization combines multiple elementary operations into a single, more complex operation. For example, a sequence of data loading, transformation, and a cryptographic operation might be fused into one optimized kernel. This significantly reduces overhead associated with intermediate data transfers between memory and compute units, as well as the overhead of scheduling individual, fine-grained operators. By reducing these "bookkeeping" costs, the overall execution time is dramatically cut.
- Load-on-Use Memory Management: This is a particularly clever optimization that addresses a common inefficiency in GPU-based acceleration. Traditional GPU pipelines often load all necessary data at the very beginning of a computation, even if certain parts are only needed much later. The speaker provided an example: in a multi-step cryptographic process, a counter might be loaded at "step 1" but only actually utilized at "step 4." During the intermediate steps, this counter occupies valuable GPU memory. If the GPU needs to perform other computations requiring different data, it might engage in costly memory swapping, moving the counter data off the GPU and then back on when it's finally needed. TensorCrypt's load-on-use strategy ensures that data, such as a counter, is only loaded onto the GPU's memory when it is actually required for an upcoming computation. This minimizes unnecessary memory occupation and reduces memory swapping, leading to more efficient utilization of GPU resources and faster execution.
Performance Evaluation and Deployment
▶ Watch: Key optimizations: operator fusion and load-on-use management (7:15)
The TensorCrypt framework was subjected to extensive evaluation to validate its effectiveness, efficiency, and broad deployability. The researchers selected three prominent symmetric ciphers for their benchmarks: AES, Chacha, and Salsa. Notably, AES and Chacha are the only two block ciphers chosen for TLS 1.3 for high-volume data encryption and decryption, underscoring their real-world relevance.
For baselines, the team used highly optimized, publicly available GPU-based implementations of these ciphers. These baselines were verified for correctness using test vectors and were chosen because they represent the most effective and popular existing acceleration methods, having been proven to outperform even dedicated CPU instruction sets like AES-NI in many scenarios.
The evaluation yielded impressive results regarding effectiveness: TensorCrypt-transformed models achieved performance gains of up to 5.4 times faster than the optimized GPU-based baselines. The paper provides a detailed breakdown of how the proposed optimizations, such as operator fusions and load-on-use memory management, contributed to these speedups, demonstrating their individual and combined impact on reducing overheads.
A key advantage highlighted by the speaker is the unparalleled flexibility in deployment. Leveraging the inherent design of neural networks, TensorCrypt models can be executed across a remarkably diverse range of hardware and software environments:
- Specialized AI Accelerators: The models were deployed and tested on Nvidia GPUs (common in data centers and high-performance computing) and Google TPUs (Tensor Processing Units, Google's custom ASICs for machine learning).
- Cloud Environments: Demonstrating scalability and accessibility, TensorCrypt models were run on cloud platforms, specifically mentioning Google Cloud.
- General-Purpose Computing: The framework's versatility extends to everyday devices, with successful deployment on desktop computers and laptops.
- Edge and IoT Devices: Crucially, TensorCrypt also showed efficacy on smartphones and IoT devices, environments where resource constraints often make strong encryption challenging. This broad compatibility is a significant departure from hardware-specific solutions and opens doors for ubiquitous, high-performance security.
The evaluation also observed a "symmetric efficiency" for both encryption and decryption operations, indicating that the performance gains are consistent regardless of the cryptographic mode (e.g., encrypting data vs. decrypting it). This comprehensive validation underscores TensorCrypt's potential to revolutionize cryptographic computation across the entire computing spectrum.
Defensive Implications
▶ Watch: Evaluation results: TensorCrypt is up to 5.4x faster (9:10)
TensorCrypt presents significant defensive implications for cybersecurity by enabling more pervasive and performant encryption across a wider range of computing environments. The ability to accelerate cryptographic computations by up to 5.4 times on existing, widely deployed neural network hardware means that organizations and individuals can implement stronger encryption without incurring prohibitive performance penalties.
Here's how TensorCrypt can bolster defensive postures:
- Ubiquitous Strong Encryption: By making encryption faster and more efficient on resource-constrained devices like IoT gadgets and smartphones, TensorCrypt facilitates the deployment of robust end-to-end encryption where it might have previously been impractical. This can significantly raise the bar for attackers attempting to compromise data at the edge or in mobile communications.
- Enhanced Performance for Privacy-Preserving Technologies: Many advanced privacy-enhancing technologies, such as homomorphic encryption, zero-knowledge proofs, and secure multi-party computation, rely heavily on complex and computationally intensive cryptographic operations. TensorCrypt's acceleration capabilities could make these technologies more practical for real-world deployment, allowing organizations to process sensitive data while maintaining privacy and compliance.
- Reduced Attack Surface for Side Channels: While the speaker acknowledged that side-channel attacks (like timing attacks) are still a concern for any implementation, the formal security proof for the transformation ensures that the design itself does not introduce new vulnerabilities. By standardizing cryptographic implementations within a neural network framework, it might be possible in future work to develop standardized, hardened implementations that are less susceptible to certain classes of side-channel attacks compared to bespoke, hand-optimized assembly code.
- Leveraging Existing Infrastructure: Defenders can leverage their existing investments in AI hardware (GPUs, TPUs) and software stacks to enhance security, rather than requiring dedicated cryptographic acceleration hardware. This reduces cost and complexity, making advanced cryptographic performance more accessible.
- Faster Patching and Updates: The framework's flexibility could potentially streamline the process of updating cryptographic primitives or patching vulnerabilities. If cryptographic algorithms are represented as graphs, changes can be propagated and recompiled efficiently across diverse platforms.
In essence, TensorCrypt empowers defenders by breaking down performance barriers, enabling the widespread adoption of robust, modern cryptography, and potentially accelerating the deployment of advanced privacy-preserving technologies that were once too slow for practical use.
Key Takeaways
- Novel Repurposing: TensorCrypt is the first framework to repurpose neural networks for efficient cryptographic computation by directly transforming cryptographic algorithms into neural network computational graphs, bypassing traditional machine learning training.
- Significant Performance Gains: The framework achieves up to 5.4 times faster cryptographic computation compared to highly optimized GPU-based baselines for ciphers like AES, Chacha, and Salsa.
- Formal Security Guarantee: TensorCrypt includes a formal proof ensuring that the transformation maintains the original cryptographic algorithm's semantic and security level, introducing no new vulnerabilities.
- Optimized for Efficiency: Key optimizations like operator fusions and load-on-use memory management significantly reduce overheads and improve memory utilization on accelerators.
- Ubiquitous Deployment: The transformed models can be deployed on a wide array of hardware (Nvidia GPUs, Google TPUs, CPUs) and platforms (cloud, desktop, laptop, smartphone, IoT devices), leveraging the diverse neural network ecosystem.
- Enables Stronger Security: By making cryptographic operations faster and more accessible, TensorCrypt facilitates the wider adoption of robust encryption, especially on resource-constrained devices, and could accelerate the practical deployment of privacy-enhancing technologies.
About the Speaker(s)
Xin Jin presented the paper "TensorCrypt: Repurposing Neural Networks for Efficient Cryptographic Computation." The work was a joint effort with collaborators and professors, as acknowledged in the introduction. Xin Jin's research focuses on finding innovative ways to optimize and accelerate fundamental computations, particularly in the realm of cryptography, by drawing parallels between established computational paradigms and emerging technologies like neural networks.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
TensorCrypt is a genuinely novel systems paper that earns its place at NDSS: the core insight — treat crypto algorithms as computational graphs and compile them onto NN accelerator stacks without any training — is non-obvious and technically substantive. 5.4x over optimized GPU baselines for TLS 1.3 ciphers is a real number, and the formal security equivalence proof shows the authors understood they were playing with fire near the crypto/ML boundary. Not a 5-star because the implementation side-channel story is left largely open and the deployment diversity claim leans more on the NN ecosystem's existing work than on TensorCrypt-specific innovation.
Heather Calloway (CISO) — WEAK
Technically credible work with a genuinely interesting insight — repurposing AI hardware stacks for cryptographic acceleration without training is clever and the 5.4x speedup headline is real. But this talk never leaves the lab. It delivers no actionable guidance for operators, security architects, or executives, and the 'defensive implications' section reads like speculative marketing copy.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025