PropertyGPT: LLM-driven Formal Verification of Smart Contracts through Retrieval-Augmented Property Generation

Ye Liu (Postdoc Researcher · Singapore Management University)

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Blockchain Security 2

Overview

The rapid proliferation of smart contracts on blockchain systems has ushered in a new era of decentralized applications, managing billions in digital assets. However, this innovation comes with significant security risks, as smart contracts are frequently targeted by attackers, leading to substantial financial losses—reportedly $38.7 million in 2024 alone. Traditional formal verification techniques, while powerful in proving program correctness and uncovering deep bugs, often struggle with the unique challenges of smart contract auditing. These methods are typically customized for common vulnerabilities or demand extensive manual input from users, hindering their scalability and adaptability to application-specific security flaws.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction: Smart contract vulnerabilities and formal methods limitations
  2. 1:53 Proposed LLM-driven formal verification pipeline with RAG
  3. 2:50 PropertyGPT core steps: retrieval, generation, and challenges
  4. 4:00 Custom property language and similarity-based scoring algorithm
  5. 5:30 Formal prover integration and diverse property dataset
  6. 6:40 Evaluation results: property generation and vulnerability detection
  7. 8:00 Generalization, top-k impact, and GPT-4 for property fixing
  8. 9:00 Real-world bug discovery and key takeaways for automation

PropertyGPT: LLM-driven Formal Verification of Smart Contracts through Retrieval-Augmented Property Generation

Speakers: Ye Liu, Postdoc Researcher, Singapore Management University

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=FOSwJGC3bO0

Overview

The rapid proliferation of smart contracts on blockchain systems has ushered in a new era of decentralized applications, managing billions in digital assets. However, this innovation comes with significant security risks, as smart contracts are frequently targeted by attackers, leading to substantial financial losses—reportedly $38.7 million in 2024 alone. Traditional formal verification techniques, while powerful in proving program correctness and uncovering deep bugs, often struggle with the unique challenges of smart contract auditing. These methods are typically customized for common vulnerabilities or demand extensive manual input from users, hindering their scalability and adaptability to application-specific security flaws.

In response to this critical challenge, Ye Liu, a Postdoc Researcher from Singapore Management University, presented "PropertyGPT: LLM-driven Formal Verification of Smart Contracts through Retrieval-Augmented Property Generation" at the NDSS Symposium. This talk introduces an innovative framework that leverages the advanced reasoning and understanding capabilities of Large Language Models (LLMs) to automate the generation of formal properties for smart contracts. By integrating LLMs into the formal verification pipeline, PropertyGPT aims to overcome the limitations of existing approaches, enabling more precise and automated detection of vulnerabilities, including those specific to a contract's unique logic.

The significance of PropertyGPT lies in its potential to revolutionize smart contract security. By automating the previously labor-intensive process of property generation, it makes advanced formal verification techniques more accessible and efficient. This work represents a crucial step towards building more secure and resilient blockchain ecosystems, protecting valuable digital assets from sophisticated attacks. The collaboration with researchers from the Hong Kong University of Science and Technology, Nanyang Technological University, and MetaTrust Labs underscores the interdisciplinary effort to tackle one of the most pressing issues in Web3 security.

Background

▶ Watch: Introduction: Smart contract vulnerabilities and formal methods limitations (0:00)

Smart contracts, as immutable programs managing high-value digital assets on public blockchains, are inherently attractive targets for malicious actors. The consequences of vulnerabilities in these contracts are often catastrophic, leading to irreversible financial losses, as evidenced by the substantial sums lost to exploits annually. While formal verification is widely recognized as the most rigorous method for ensuring software correctness and uncovering subtle, hard-to-find bugs beyond the scope of traditional testing, its application to smart contracts faces significant hurdles.

Existing formal methods for smart contracts, such as symbolic execution, SMT solving, and model checking, have made notable contributions. However, many of these approaches are either highly specialized, designed to detect only a narrow range of common vulnerabilities like re-entrancy, or they demand considerable manual effort. General-purpose tools often require users to provide detailed specifications, invariants, or pre/post-conditions—known as formal properties—which are notoriously difficult and time-consuming to write correctly, especially for complex or novel contract logic. This high barrier to entry limits the widespread adoption and full automation of formal verification, leaving a significant gap in the ability to identify application-specific vulnerabilities that deviate from common patterns.

The core problem, therefore, is how to automatically adapt formal verification techniques to the diverse and evolving landscape of smart contracts, enabling the detection of vulnerabilities unique to each contract's specific design and functionality, without requiring intensive user input. This challenge has become particularly pertinent with the advent of powerful LLMs, which have demonstrated remarkable capabilities in program understanding, reasoning, and code generation. Inspired by these advancements, the PropertyGPT project sought to harness LLMs to bridge the gap between complex formal verification techniques and the practical need for automated, adaptable smart contract security analysis.

Key Findings

▶ Watch: PropertyGPT core steps: retrieval, generation, and challenges (2:50)

PropertyGPT's comprehensive evaluation yielded several significant findings, demonstrating its effectiveness in both property generation and vulnerability detection, alongside its practical applicability and generalization capabilities.

  1. High-Quality Property Generation: PropertyGPT achieved an impressive 80% coverage of equivalent properties found in the ground truth, based on a conservative precision calculation. The researchers noted that if correctly verified properties not explicitly in the ground truth were also counted, the actual precision would be significantly higher, indicating the LLM's capacity to generate valid, novel properties.
  1. Superior Vulnerability Detection: In a comparative analysis against four other leading security tools, PropertyGPT outperformed all competitors in detecting 13 representative CVEs, successfully identifying 9 of them. Furthermore, when evaluated on 24 attack incident projects from the Smartian benchmark, PropertyGPT accurately identified 17 incidents, showcasing its robustness against real-world exploits.
  1. Demonstrated Generalization: The study measured the code similarity between subject contracts (from attack incidents) and the reference code used for property generation (from the Storra audit projects). An average code similarity of 0.68 was observed. This moderate similarity score indicates that PropertyGPT is not merely overfitting to highly similar codebases but possesses a certain level of generalization, allowing it to apply learned patterns to contracts with varying structures and functionalities.
  1. Optimized Top-K Selection: An investigation into the impact of different top-K selections (the number of most similar reference properties considered) on generation accuracy revealed that selecting the top two most similar properties provided an optimal balance between recall and precision in property generation.
  1. Effective Property Fixes with LLMs: A crucial aspect of PropertyGPT's practical utility is its ability to correct uncompilable generated properties. Using the GPT-4 model, 84% of uncompilable properties were successfully fixed within five attempts, leading to an overall success rate of 87%. This highlights the LLM's capability not only to generate but also to iteratively refine and correct formal specifications.
  1. Real-World Impact and Bug Bounties: Beyond benchmark results, PropertyGPT successfully identified genuine bugs in real-world smart contracts, leading to the discovery of vulnerabilities and the receipt of bug bounty rewards. This tangible outcome underscores the framework's practical value and its potential to contribute directly to enhancing smart contract security.

Technical Deep Dive

▶ Watch: Formal prover integration and diverse property dataset (5:30)

The core innovation of PropertyGPT lies in its LLM-driven formal verification pipeline, specifically powered by Retrieval-Augmented Property Generation (RAG). This architecture is designed to address the inherent challenges of automating formal property generation for diverse smart contracts, moving beyond predefined vulnerability patterns to detect application-specific flaws.

The PropertyGPT framework operates through a multi-step process:

  1. Knowledge Base Construction: The foundation of PropertyGPT is a robust knowledge base comprising 623 high-quality, human-written security properties collected from 23 audit projects by Storra, a renowned security company. These properties are categorized into six types (e.g., DeFi, token contracts) and are paired with their corresponding smart contract code. This diverse dataset serves as the "domain knowledge" that the LLM leverages.
  1. Retrieval of Reference Code and Properties: Given a new subject code (the smart contract to be verified), PropertyGPT first queries its knowledge base to identify the most similar reference code. This retrieval is crucial for contextualizing the property generation process. Subsequently, the human-written properties associated with this similar reference code are retrieved, serving as reference properties for the LLM.
  1. LLM-driven Property Generation: For each retrieved reference property, an LLM is employed to generate a new, tailored property for the subject code. This generation is guided by a carefully designed prompt (the specifics of which are detailed in the research paper). Initially, the researchers attempted to use Storra's existing property language but found that LLMs struggled to generate compilable and complete properties. The difficulty stemmed from the LLM's challenge in reasoning effectively across two fundamentally different language paradigms (natural language for instructions and a highly specialized formal property language).
  1. Custom Property Language: To overcome this limitation, PropertyGPT introduces its own property language. This custom language is designed to be highly similar to Solidity, the popular smart contract programming language. It supports three types of properties:
  • Contract-level invariants: Properties that must hold true throughout the contract's lifecycle.
  • Function-level pre and post-conditions: Properties that define the state before and after a function's execution.
  • User-defined rules: Flexible assertions and assumptions based on specific security concerns.

This design choice proved critical, as the LLM could generate compilable properties much more easily within a language structure closely resembling the source code it was analyzing.

  1. Scoring and Selection Algorithm: The LLM may generate multiple candidate properties. To select the most suitable ones, PropertyGPT employs a sophisticated scoring algorithm. This algorithm evaluates candidates based on four distinct similarity metrics:
  • Similarity between the subject code and the reference code.
  • Similarity between the reference property and the newly generated property.
  • A summary-based metric is also introduced to capture functional similarities between different code segments, acknowledging that functionally identical code can have varied syntactic forms.

The algorithm utilizes a linear model trained to compute optimal weight coefficients for these metrics, ensuring that the selected properties are highly relevant and accurate. The training of these coefficients represents the primary computational cost of the PropertyGPT pipeline, though inference (property generation) is noted as being very cheap, especially with open-source LLMs.

  1. Formal Verification with SoC: The selected properties are then passed to PropertyGPT's formal prover. This prover is built upon SoC (Source of Contracts), a previous work by the authors that functions as a source-level symbolic execution engine for Solidity. The prover performs source code-level analysis, collecting all symbolic paths through the contract. During this phase, it checks whether the generated properties are satisfied under various execution scenarios. If a property violation is detected, it implies a potential vulnerability in the smart contract. Due to the lack of ground truth for many detected issues, identified violations are subsequently confirmed through manual code review to ascertain the presence of a genuine vulnerability.

This integrated approach, combining LLM intelligence with established formal verification techniques and a tailored property language, allows PropertyGPT to automatically generate precise, application-specific formal properties, significantly advancing the automation and effectiveness of smart contract security auditing.

Demo / Proof of Concept

▶ Watch: Evaluation results: property generation and vulnerability detection (6:40)

While the talk did not feature a live, interactive demonstration, the comprehensive evaluation presented serves as a robust proof of concept for PropertyGPT's methodology and its effectiveness in real-world scenarios. The researchers meticulously evaluated the framework across multiple dimensions, showcasing its capabilities in both property generation and vulnerability detection.

The initial proof of concept for property generation involved testing PropertyGPT's ability to produce relevant formal specifications. Using 19 properties from nine Storra audit projects as test cases, with the remaining 604 properties serving as the knowledge base, PropertyGPT demonstrated its capacity to generate properties equivalent to 80% of the ground truth. This high coverage, even with a conservative precision calculation, effectively validates the LLM-driven Retrieval-Augmented Property Generation mechanism.

For vulnerability detection, the proof of concept involved two key benchmarks:

  1. CVE Detection: PropertyGPT was benchmarked against four other leading security tools in detecting 13 representative CVEs (Common Vulnerabilities and Exposures). The results showed PropertyGPT outperforming all other tools by successfully identifying 9 of these 13 CVEs. This comparative performance highlights the efficacy of the generated properties in uncovering known, critical vulnerabilities.
  2. Smartian Attack Incidents: Further validation came from evaluating PropertyGPT on 24 attack incident projects from the Smartian benchmark, a collection of real-world exploited smart contracts. PropertyGPT successfully identified 17 of these 24 attack incidents, demonstrating its ability to detect vulnerabilities that have led to actual exploits.

A crucial aspect of the proof of concept was also to assess the framework's generalization capacity. By measuring the average code similarity between the subject code from the Smartian projects and the reference code from the Storra projects, an average similarity of 0.68 was observed. This intermediate similarity score is significant because it indicates that PropertyGPT is not merely identifying vulnerabilities in highly similar code but can generalize its property generation to contracts with differing structures, thus proving its adaptability beyond exact matches.

Finally, the most compelling proof of concept came from PropertyGPT's ability to discover real-world bugs in actual smart contracts, leading to the award of bug bounties. This practical outcome unequivocally validates the framework's utility and impact in enhancing the security posture of live blockchain applications, moving beyond theoretical benchmarks to tangible security improvements.

Defensive Implications

▶ Watch: Real-world bug discovery and key takeaways for automation (9:00)

PropertyGPT introduces several critical defensive implications for smart contract security, offering new strategies for developers, auditors, and the broader Web3 ecosystem.

  1. Automated, Application-Specific Verification: For smart contract developers and security auditors, PropertyGPT offers a powerful mechanism to move beyond generic vulnerability scanners. By automatically generating tailored formal properties, it enables a deeper, application-specific verification of contract logic. This means defenders can identify vulnerabilities unique to a contract's business rules or complex interactions, rather than being limited to common patterns like re-entrancy or integer overflows. Integrating such a tool into continuous integration/continuous deployment (CI/CD) pipelines could allow for proactive security checks throughout the development lifecycle.
  1. Reduced Manual Effort in Formal Verification: Formal verification has traditionally been the domain of highly specialized experts, largely due to the immense effort required to manually write formal properties. PropertyGPT significantly lowers this barrier. It automates the most laborious step, allowing formal verification experts to focus their valuable time on analyzing the generated properties, refining complex specifications, or investigating the root causes of reported violations, rather than drafting properties from scratch. This could democratize access to advanced security analysis.
  1. Importance of Domain-Specific Property Languages: The finding that LLMs perform significantly better with a Solidity-like property language is a key insight for tool developers. It suggests that future LLM-powered security tools should prioritize designing their intermediate representations or property languages to closely mirror the target programming language. This reduces the cognitive load on the LLM, leading to more accurate, compilable, and useful outputs.
  1. Enhanced Understanding of LLM Capabilities and Limitations: The discussion around LLM "hallucination" and the speaker's perspective that generating properties "not similar as the ground truth" can sometimes be a positive (offering novel, valid perspectives) is crucial. Defenders should view LLM outputs as intelligent suggestions or candidates, requiring expert review rather than blind trust. This iterative human-in-the-loop approach combines the LLM's scale with human expertise.
  1. Future Potential for Optimizing Existing Tools: The talk points to the future potential of LLMs to optimize existing formal techniques. For instance, LLMs could be used to mitigate the state explosion problem in symbolic execution, a major bottleneck for analyzing complex smart contracts. By intelligently guiding the symbolic execution engine, LLMs could help prioritize paths or prune irrelevant states, making these powerful but computationally intensive tools more scalable and practical.

In summary, PropertyGPT provides a blueprint for a new generation of security tools that leverage LLMs to make advanced formal verification more automated, efficient, and adaptable. Defenders can use these insights to build more robust security practices, integrate intelligent property generation into their workflows, and better protect the valuable assets managed by smart contracts.

Key Takeaways

  • LLMs Automate Formal Property Generation: PropertyGPT successfully demonstrates that Large Language Models can effectively generate formal security properties for smart contracts, significantly automating a previously manual and labor-intensive step in formal verification.
  • Retrieval-Augmented Generation (RAG) is Key: The framework's reliance on high-quality human-written properties and retrieval-augmented generation is crucial for grounding LLM outputs in established security domain knowledge, ensuring relevance and accuracy.
  • Custom Property Language Enhances LLM Performance: Developing a custom property language that closely resembles Solidity (the target programming language) dramatically improves the LLM's ability to generate compilable, complete, and correct formal specifications, overcoming challenges with existing formal languages.
  • Superior Vulnerability Detection: PropertyGPT outperforms traditional security tools in detecting diverse smart contract vulnerabilities, identifying 9 out of 13 representative CVEs and 17 out of 24 Smartian attack incidents, showcasing its practical effectiveness.
  • Practical Generalization and Real-World Impact: The framework exhibits a degree of generalization (average code similarity of 0.68) and has successfully found real-world bugs, leading to bug bounties, proving its applicability beyond specific code patterns.
  • Future Directions for LLM-Enhanced Security: Future research can explore using LLMs to optimize existing formal techniques (e.g., mitigating state explosion in symbolic execution) and investigate the soundness and completeness of LLM reasoning in formal tasks.

About the Speaker(s)

The primary speaker for "PropertyGPT: LLM-driven Formal Verification of Smart Contracts through Retrieval-Augmented Property Generation" was Ye Liu, currently a Postdoc Researcher at Singapore Management University.

The work presented is a collaborative effort, involving authors from several esteemed institutions: the Hong Kong University of Science and Technology, Nanyang Technological University, and MetaTrust Labs in Singapore. This interdisciplinary collaboration highlights a concerted academic and industry effort to advance the state of smart contract security.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Competent systems research that combines RAG with formal verification for smart contract property generation — a real problem, a real contribution, and actual bug bounties to show for it. The work is legitimate but not groundbreaking; the core insight (LLMs are better at generating specs in a Solidity-like DSL than in an alien formal language) is sensible engineering rather than a paradigm shift, and the evaluation numbers, while solid, leave meaningful gaps unaddressed.

Heather Calloway (CISO) — PASS

Technically credible academic research on LLM-assisted formal verification for smart contracts. Outside my lane — no governance angle, no institutional accountability dimension, and no relevance to the security programs I run or advise.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025