In a report with academic partners, OpenAI documented eight real deployments where coding agents modernized neglected research software — including a 60x speedup in RNA-sequencing quality control and a 20,000-line C++ genome aligner rewritten in Rust at 99.8% parity. Five projects used Codex alone; three combined it with Anthropic's Claude Code. The report is blunt about the limit: agents finished ambitious engineering tasks but could not judge whether results were scientifically correct, sometimes presenting flawed code with full confidence.