OpenAI published MentalHealthBench on Wednesday, an open test of how models handle mental health conversations. More than 80 psychologists and psychiatrists in 22 countries wrote 1,215 conversations and 5,262 scoring rules. GPT-6 Astra led with 57.3 percent, up from 32.1 percent for a 2025 model.