Tools
AI Minute Newsroom
2026-08-24
Developers noticed the effort dial reading low. Anthropic says it is testing serving configs, and the number means something different this week.
Over the weekend of 22 August, Claude Code users posted that the tool's "effort" indicator was showing unexpectedly low values — one screenshot reported 10 out of 100 — and that identical tasks were behaving very differently between Opus 5 and an older Opus 4.6, in one case 43 minutes against 2 minutes for the same job. The thread reached the front page of Hacker News. Thariq, who leads the Claude Code team at Anthropic, replied in public that the company is currently testing API serving configurations and that during that test the numerical effort value is "mapped differently" than usual. He said the effort a user selects is the effort they get, and that the company's own evaluations show no drop in performance. The rest of the thread is not so easily settled: a long run of users describing Opus 5 output as more verbose and more elaborate than they want, some switching back to 4.6 or to rival tools. Those are impressions, not measurements, and impressions of model quality are famously unreliable — but the specific thing Anthropic confirmed is that a number shown in the interface did not mean this week what it meant last week.
Why it mattersThis is a small incident with a large lesson about buying capability by the token. What you rent from a model provider is not a fixed artefact: the weights may be stable, but the serving configuration around them — quantisation, routing, effort mapping, context handling — is tuned continuously and usually without a release note. Here the change was visible because it moved a number on screen; most such changes are not. If your work depends on model behaviour staying constant, the practical response is to keep your own small evaluation set and run it periodically, rather than relying on the vendor's dashboard or on your own sense that something feels off. And treat "the model got worse" posts with care: they are sometimes right, and they are almost never evidence.
✓ Verified · 2 sources
Read in the app — free, in 9 languages
Related stories
A 300-million-parameter model guesses what the big one will say next — and a MacBook answers 2.87 times faster
2026-08-23Asked how many of their 'agents' actually finish a multi-step job unattended, seven in ten companies said a quarter or fewer
2026-08-23A search index built for coding agents rather than people: 70 million READMEs, issues and pull requests, and no key needed to try it
2026-08-23The standard that plugs AI into your tools was built for a human clicking 'allow'. Its new roadmap admits the caller is now a machine.
2026-08-23A 27-billion-parameter model beat Opus 4.8 and GPT-5.5 at reproducing published papers
2026-08-23