aiminute. ← All AI news
Research 2026-08-15

Asked to write code and prove it correct at the same time, the best agent managed 27 out of 43 — and zero on the hard ones

Asked to write code and prove it correct at the same time, the best agent managed 27 out of 43 — and zero on the hard ones

A team including Dawn Song published Vero on arXiv on 13 August, described as the first benchmark for repository-level verified code synthesis. Instead of asking an agent for a function that passes tests, Vero asks it to produce a working implementation across a multi-module codebase and a machine-checkable proof that the implementation matches a formal specification. It contains 43 instances built from real repositories, spanning proof and programming languages including Lean 4, Dafny, Verus and Coq, and domains from cryptographic protocols to distributed systems. The strongest agent configuration the authors tested fully solved 27 of the 43, and closed no specifications at all on the hardest repositories.

Why it mattersAlmost every claim you read about agents writing software rests on tests passing, and tests only cover the cases someone thought to write. A machine-checked proof is the one form of evidence that does not care how clever the model sounded — either the proof compiles or it does not. This is the first time anyone has measured agents against that standard at the scale of a whole repository, and the answer is: they clear the easy two-thirds and do not touch the hard third. That is a useful number to hold on to the next time a coding agent is described as reliable.
#Coding#AI Agents

✓ Verified · 1 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Machine learning read the shape of sick brain cells and picked out nine already-approved drugs that calmed them down
2026-08-21
Rumour: the anonymous model that just topped a coding benchmark, for free, is said to be Zhipu's unreleased flagship
2026-08-21
DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21
Nvidia is paying $6 billion for the machine that builds a rival's models — and hiring 109 of the people who ran it
2026-08-21
Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21