Specific Labs published Real-SWE, a test built from private production code licensed from real companies. Eight frontier models solved between 16.2 and 38.8 percent of the tasks, with Claude Fable 5.1 on top. GPT-6 Astra reached 33.8 percent and Gemini 3.8 Flash 31.2 percent on the same set.