Research
AI Minute Newsroom
2026-08-15
Asked to write code and prove it correct at the same time, the best agent managed 27 out of 43 — and zero on the hard ones
A team including Dawn Song published Vero on arXiv on 13 August, described as the first benchmark for repository-level verified code synthesis. Instead of asking an agent for a function that passes tests, Vero asks it to produce a working implementation across a multi-module codebase and a machine-checkable proof that the implementation matches a formal specification. It contains 43 instances built from real repositories, spanning proof and programming languages including Lean 4, Dafny, Verus and Coq, and domains from cryptographic protocols to distributed systems. The strongest agent configuration the authors tested fully solved 27 of the 43, and closed no specifications at all on the hardest repositories.
Why it mattersAlmost every claim you read about agents writing software rests on tests passing, and tests only cover the cases someone thought to write. A machine-checked proof is the one form of evidence that does not care how clever the model sounded — either the proof compiles or it does not. This is the first time anyone has measured agents against that standard at the scale of a whole repository, and the answer is: they clear the easy two-thirds and do not touch the hard third. That is a useful number to hold on to the next time a coding agent is described as reliable.
✓ Verified · 1 sources
Read in the app — free, in 9 languages
Related stories
A model learned to write without the method that trains every AI.
2026-10-06TikTok put a shopping chatbot inside the video you are watching.
2026-10-06Reflection will hand out a 501-billion-parameter model for free.
2026-10-06Wikipedia's owner says OpenAI agents may have caused a May outage.
2026-10-06Meta's assistant keeps an hourly file on everyone in your life.
2026-10-05