AI Post Transformers

From Natural Language to Verified Dafny Code


Listen Later

This episode explores a 2026 study on turning long natural-language programming problems into Dafny code that can be formally verified, asking whether AI systems can produce code that is not just fluent but provably correct. It explains how Dafny uses preconditions, postconditions, loop invariants, and proof obligations, and why weak specifications can lead to vacuous “verified” programs that still fail to capture the real task. The discussion highlights the paper’s NL2VC-60 benchmark of hand-written verified solutions to UVa-style algorithm problems, along with experiments comparing plain prompting, signature-guided prompting, and self-healing loops that revise code using verifier feedback and additional uDebug testing. Listeners would find it interesting because it gets at the core trust problem in AI coding: whether formal methods can make generated software more reliable, and where the real bottleneck remains the human effort required to write strong specifications.
Sources:
1. From Natural Language to Verified Code: Toward AI Assisted Problem-to-Code Generation with Dafny-Based Formal Verification — Md Erfan, Md Kamal Hossain Chowdhury, Ahmed Ryan, Md Rayhanur Rahman, 2026
http://arxiv.org/abs/2604.22601
2. Dafny: An Automatic Program Verifier for Functional Correctness — K. Rustan M. Leino, 2010
https://scholar.google.com/scholar?q=Dafny%3A+An+Automatic+Program+Verifier+for+Functional+Correctness
3. seL4: Formal Verification of an Operating-System Kernel — Gerwin Klein, Kevin Elphinstone, Gernot Heiser, June Andronick, David Cock, et al., 2009
https://scholar.google.com/scholar?q=seL4%3A+Formal+Verification+of+an+Operating-System+Kernel
4. Formal verification of a realistic compiler — Xavier Leroy, 2009
https://scholar.google.com/scholar?q=Formal+verification+of+a+realistic+compiler
5. Modularity, Code Specialization, and Zero-Cost Abstractions for Program Verification — Son Ho, Aymeric Fromherz, Jonathan Protzenko, 2021
https://scholar.google.com/scholar?q=Modularity%2C+Code+Specialization%2C+and+Zero-Cost+Abstractions+for+Program+Verification
6. Towards AI-Assisted Synthesis of Verified Dafny Methods — Md Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James Noble, 2024
https://scholar.google.com/scholar?q=Towards+AI-Assisted+Synthesis+of+Verified+Dafny+Methods
7. DafnyBench: A Benchmark for Formal Software Verification — Chloe Loughridge et al., 2024
https://scholar.google.com/scholar?q=DafnyBench%3A+A+Benchmark+for+Formal+Software+Verification
8. Can LLMs Enable Verification in Mainstream Programming? — Aleksandr Shefer, Igor Engel, Stanislav Alekseev, Daniil Berezun, Ekaterina Verbitskaia, Anton Podkopaev, 2025
https://scholar.google.com/scholar?q=Can+LLMs+Enable+Verification+in+Mainstream+Programming%3F
9. Dafny as Verification-Aware Intermediate Language for Code Generation — Yue Chen Li, Stefan Zetzsche, Siva Somayyajula, 2025
https://scholar.google.com/scholar?q=Dafny+as+Verification-Aware+Intermediate+Language+for+Code+Generation
10. ATLAS: Automated Toolkit for Large-Scale Verified Code Synthesis — Mantas Baksys et al., 2025
https://scholar.google.com/scholar?q=ATLAS%3A+Automated+Toolkit+for+Large-Scale+Verified+Code+Synthesis
11. DafnyPro: LLM-Assisted Automated Verification for Dafny Programs — Debangshu Banerjee, Olivier Bouissou, Stefan Zetzsche, 2026
https://scholar.google.com/scholar?q=DafnyPro%3A+LLM-Assisted+Automated+Verification+for+Dafny+Programs
12. Neuro Symbolic Reasoning for Planning: Counterexample Guided Inductive Synthesis using Large Language Models and Satisfiability Solving — Sumit Kumar Jha et al., 2023
https://scholar.google.com/scholar?q=Neuro+Symbolic+Reasoning+for+Planning%3A+Counterexample+Guided+Inductive+Synthesis+using+Large+Language+Models+and+Satisfiability+Solving
13. Property-Guided LLM Program Synthesis for Planning — Andre G. Pereira, Augusto B. Correa, Jendrik Seipp, 2026
https://scholar.google.com/scholar?q=Property-Guided+LLM+Program+Synthesis+for+Planning
14. Finding Inductive Loop Invariants using Large Language Models — Adharsh Kamath et al., 2023
https://scholar.google.com/scholar?q=Finding+Inductive+Loop+Invariants+using+Large+Language+Models
15. LLM For Loop Invariant Generation and Fixing: How Far Are We? — Mostafijur Rahman Akhond, Saikat Chakraborty, Gias Uddin, 2025
https://scholar.google.com/scholar?q=LLM+For+Loop+Invariant+Generation+and+Fixing%3A+How+Far+Are+We%3F
16. Type-Constrained Code Generation with Language Models — Niels Mundler et al., 2025
https://scholar.google.com/scholar?q=Type-Constrained+Code+Generation+with+Language+Models
17. Invariant-based Program Repair — Omar I. Al-Bataineh, 2024
https://scholar.google.com/scholar?q=Invariant-based+Program+Repair
18. AI Post Transformers: Program Synthesis with Large Language Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-program-synthesis-with-large-language-mo-b962ec.mp3
19. AI Post Transformers: SGLang for Faster Structured LLM Programs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-sglang-for-faster-structured-llm-program-c59f1c.mp3
20. AI Post Transformers: SkillsBench for Evaluating Agent Skills — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-skillsbench-for-evaluating-agent-skills-58bb1e.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof