AI Post Transformers

DafnyBench and LLMs for Formal Verification


Listen Later

This episode explores DafnyBench, a benchmark for testing whether large language models can help with one of formal verification’s hardest practical bottlenecks: reconstructing the missing assertions and loop invariants that make Dafny programs verifiable. It explains how formal verification differs from ordinary testing and from theorem proving, and why the paper deliberately frames the task as restoring proof hints in existing verified programs rather than synthesizing correct software from scratch. The discussion digs into benchmark design, including the dataset of 782 single-file Dafny programs, the rule that models must infer both the content and placement of missing hints, and the importance of excluding shortcut tricks like disabling verification. It also highlights a crucial result nuance: 208 files already verify after hint removal, so the reported top score of about 67.8% is more informative when translated into genuine recovery performance on the subset that actually needs new annotations.
Sources:
1. DafnyBench: A Benchmark for Formal Software Verification — Chloe Loughridge, Qinyi Sun, Seth Ahrenbach, Federico Cassano, Chuyue Sun, Ying Sheng, Anish Mudide, Md Rakib Hossain Misu, Nada Amin, Max Tegmark, 2024
http://arxiv.org/abs/2406.08467
2. Clover: Closed-Loop Verifiable Code Generation — Chuyue Sun, Ying Sheng, Oded Padon, Clark Barrett, 2024
https://scholar.google.com/scholar?q=Clover%3A+Closed-Loop+Verifiable+Code+Generation
3. Towards AI-Assisted Synthesis of Verified Dafny Methods — Md Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James Noble, 2024
https://scholar.google.com/scholar?q=Towards+AI-Assisted+Synthesis+of+Verified+Dafny+Methods
4. LeanDojo: Theorem Proving with Retrieval-Augmented Language Models — Kaiyu Yang et al., 2023
https://scholar.google.com/scholar?q=LeanDojo%3A+Theorem+Proving+with+Retrieval-Augmented+Language+Models
5. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code — Naman Jain et al., 2024
https://scholar.google.com/scholar?q=LiveCodeBench%3A+Holistic+and+Contamination+Free+Evaluation+of+Large+Language+Models+for+Code
6. Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification — Xu Xu et al., 2025
https://scholar.google.com/scholar?q=Local+Success+Does+Not+Compose%3A+Benchmarking+Large+Language+Models+for+Compositional+Formal+Verification
7. A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification — Norbert Tihanyi et al., 2023
https://scholar.google.com/scholar?q=A+New+Era+in+Software+Security%3A+Towards+Self-Healing+Software+via+Large+Language+Models+and+Formal+Verification
8. AI Post Transformers: LLM Agents Reason About Code Without Running It — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-15-llm-agents-reason-about-code-without-run-2a1876.mp3
9. AI Post Transformers: SkillsBench for Evaluating Agent Skills — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-skillsbench-for-evaluating-agent-skills-58bb1e.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof