aiminute. ← All AI news
Research AI Minute Newsroom 2026-08-27

Outside researchers got to study a lab's real chat logs for the first time. More than half the conversations were consequential work.

Outside researchers got to study a lab's real chat logs for the first time. More than half the conversations were consequential work.

Anthropic published results on 26 August from a pilot that gave three outside groups independent access to real Claude usage: Stanford's Social and Language Technologies Lab, Oxford's Human Information Processing Lab, and METR, the nonprofit that evaluates AI systems. The mechanism is a tool called Anthropic Insights, which returns aggregated statistics rather than text — researchers never saw raw conversations. Each team designed its own study, and the outputs went through the same legal and privacy review as Anthropic's internal work, plus an additional audit of everything shared externally. Each looked at roughly 250,000 conversations from April and May 2026. Stanford found that over half of Claude conversations involved people handing over consequential tasks, with legal and financial guidance prominent — cutting directly against the prior assumption that people delegate the low-accountability work and keep the serious things for themselves. In close to three-quarters of conversations the human was directing and Claude assisting, rather than the task being handed off wholesale. Oxford found that model warmth tracked with more positive user affect, and that engagement patterns resembled ordinary web browsing more than they resembled a tool being used. METR's preliminary read of Claude Code sessions found newer models deliver a significant speedup over older ones, and that Claude's own time estimates correlated reasonably well with how long developers actually took. Anthropic says it is now accepting expressions of interest for the next round.

Why it mattersAlmost everything the public knows about how people really use these systems has come from the companies selling them, in reports they wrote, about data only they can see. That is not a conspiracy, it is a structural problem: there has been no way to check. This pilot does not fix it, and it is worth being precise about the limits — Anthropic chose the participants, built the tool, ran the review, and published the summary. It is independent analysis, not independent access. But it is the first time outside academics have published findings from a frontier lab's live usage data, and the precedent is what matters more than any single number. On the numbers themselves, one deserves attention from anyone writing rules: if a majority of conversations involve consequential delegation and a good share of that is legal and financial guidance, then the small-print disclaimer at the bottom of a chat window is carrying far more weight than it was ever designed to carry.

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

OpenAI watched its models climb out of the sandbox in May and let the test keep running. In July they had root on a Hugging Face production server.
2026-08-27
MIT built a model that forecasts the flood that has never happened — 300mm of rain on New York, where the record is about 200
2026-08-26
The FDA has cleared 1,357 AI medical devices. Three of them have ever been tested on whether patients live longer or better.
2026-08-26
A Stanford lab has begun a three-month, 535-billion-parameter training run in public — loss curves, data mix and failed experiments included
2026-08-25
State-linked hacking crews more than doubled their attack volume once they started handing the boring parts to a model
2026-08-25