We are in Menlo Park this week for the first Formal Methods × AI, hosted at SRI International from September 30 to October 2 and organized by Atlas Computing. It is a small, invitation-only gathering of around fifty people. Invitees included Leo de Moura, Adam Chlipala, Clark Barrett, Swarat Chaudhuri, Stuart Russell and Max Tegmark.
Four things stood out.
Proving is still a bottleneck. Models can write Lean, but getting a proof to close on anything beyond a benchmark problem is slow, expensive and unreliable. Every working group hit the same wall.
Specifications are the other half. A proof is only as good as the statement it proves. Much of the discussion was about how to write specifications for AI-generated code, and for AI systems themselves, precisely enough to be worth proving. Nobody has a good answer yet.
There is no Mathlib for software. Mathematics has a unified, machine-checked library to build on. Software does not. CSLib is a promising start, but the gap is large, and the field has no real benchmarks yet.
The field is forming. Eight new organizations at the intersection have appeared in the last year, Cajal among them, alongside Math Inc., Theorem Labs, Axiom Math, Principia Labs and Sigil Logic. Lean dominates the conversation.
We left with a clear picture of the work: make proving fast enough that it stops being a limiting step, and build the infrastructure AI will need.