AI Post Transformers

Dafny for Trustworthy AI Code Generation


Listen Later

This episode explores a 2025 paper on using Dafny as a hidden, verification-aware intermediate language for AI code generation, where a model first produces a formal specification and verified implementation before compiling it into ordinary Python. It examines the paper’s central trust claim: formal verification can prove that the generated code satisfies the hidden spec, but it cannot prove that the hidden spec actually matches the user’s intent, making spec alignment a separate and critical failure point. The discussion uses the paper’s fibfib example and HumanEval results to unpack that distinction, noting that the Dafny-only pipeline trails direct Python generation, while the best reported score comes only after falling back to unverified Python when the verification loop fails to converge. Listeners would find it interesting because it gives a concrete, nuanced look at where AI coding assistants can become more reliable, where the guarantees stop, and why neuro-symbolic workflows may matter most for tightly specified code like algorithms, parsers, and protocol logic.
Sources:
1. Dafny as Verification-Aware Intermediate Language for Code Generation — Yue Chen Li, Stefan Zetzsche, Siva Somayyajula, 2025
http://arxiv.org/abs/2501.06283
2. Dafny: An Automatic Program Verifier for Functional Correctness — K. Rustan M. Leino, 2010
https://scholar.google.com/scholar?q=Dafny%3A+An+Automatic+Program+Verifier+for+Functional+Correctness
3. Towards AI-Assisted Synthesis of Verified Dafny Methods — Md Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James Noble, 2024
https://scholar.google.com/scholar?q=Towards+AI-Assisted+Synthesis+of+Verified+Dafny+Methods
4. Clover: Closed-Loop Verifiable Code Generation — Chuyue Sun, Ying Sheng, Oded Padon, Clark Barrett, 2024
https://scholar.google.com/scholar?q=Clover%3A+Closed-Loop+Verifiable+Code+Generation
5. VerMCTS: Synthesizing Multi-Step Programs using a Verifier, a Large Language Model, and Tree Search — David Brandfonbrener, Simon Henniger, Sibi Raja, Tarun Prasad, Chloe Loughridge, Federico Cassano, Sabrina Ruixin Hu, Jianang Yang, William E. Byrd, Robert Zinkov, Nada Amin, 2024
https://scholar.google.com/scholar?q=VerMCTS%3A+Synthesizing+Multi-Step+Programs+using+a+Verifier%2C+a+Large+Language+Model%2C+and+Tree+Search
6. Laurel: Generating Dafny Assertions Using Large Language Models — Eric Mugnier, Emmanuel Anaya Gonzalez, Ranjit Jhala, Nadia Polikarpova, Yuanyuan Zhou, 2024
https://scholar.google.com/scholar?q=Laurel%3A+Generating+Dafny+Assertions+Using+Large+Language+Models
7. DafnyBench: A Benchmark for Formal Software Verification — Chloe Loughridge et al., 2024
https://scholar.google.com/scholar?q=DafnyBench%3A+A+Benchmark+for+Formal+Software+Verification
8. Evaluating Large Language Models Trained on Code — Mark Chen et al., 2021
https://scholar.google.com/scholar?q=Evaluating+Large+Language+Models+Trained+on+Code
9. Baking for Dafny: A CakeML Backend for Dafny — Daniel Nezamabadi, Magnus Myreen, 2025
https://scholar.google.com/scholar?q=Baking+for+Dafny%3A+A+CakeML+Backend+for+Dafny
10. Intent-aligned Formal Specification Synthesis via Traceable Refinement — Zhe Ye et al., 2026
https://scholar.google.com/scholar?q=Intent-aligned+Formal+Specification+Synthesis+via+Traceable+Refinement
11. Combining LLM Code Generation with Formal Specifications and Reactive Program Synthesis — William Murphy et al., 2024
https://scholar.google.com/scholar?q=Combining+LLM+Code+Generation+with+Formal+Specifications+and+Reactive+Program+Synthesis
12. StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback — Shihan Dou et al., 2024
https://scholar.google.com/scholar?q=StepCoder%3A+Improve+Code+Generation+with+Reinforcement+Learning+from+Compiler+Feedback
13. InspectCoder: Dynamic Analysis-Enabled Self Repair through interactive LLM-Debugger Collaboration — Yunkun Wang et al., 2025
https://scholar.google.com/scholar?q=InspectCoder%3A+Dynamic+Analysis-Enabled+Self+Repair+through+interactive+LLM-Debugger+Collaboration
14. FormalSpecCpp: A Dataset of C++ Formal Specifications created using LLMs — Madhurima Chakraborty et al., 2025
https://scholar.google.com/scholar?q=FormalSpecCpp%3A+A+Dataset+of+C%2B%2B+Formal+Specifications+created+using+LLMs
15. AI Post Transformers: From Natural Language to Verified Dafny Code — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-14-from-natural-language-to-verified-dafny-8abed9.mp3
16. AI Post Transformers: Program Synthesis with Large Language Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-program-synthesis-with-large-language-mo-b962ec.mp3
17. AI Post Transformers: Generative File Systems: Replacing Code with Formal Specifications — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-generative-file-systems-replacing-code-w-414029.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof