Anthropic Opens 250,000 Claude Chats to Outside Researchers
Anthropic has granted external research labs privacy-preserving access to 250,000 real Claude conversations, establishing a new precedent for independent oversight of commercial AI usage.

Anthropic has launched a pilot program allowing three independent research groups to analyze approximately 250,000 real conversations from Claude.ai and Claude Code. The partners—the Stanford Social and Language Technologies (SALT) Lab, the Oxford Human Information Processing Lab, and METR—examined user interactions sampled from April to May 2026. To protect user privacy, researchers used a query tool called Anthropic Insights, formerly known as Clio. This system aggregates Claude's answers to researcher questions into categories, ensuring external teams never view raw chat logs. Imperial College London conducted a third-party audit to verify the privacy protections.
The initial findings challenge common assumptions about how humans interact with AI. Stanford's SALT Lab discovered that over half of the analyzed Claude conversations involved consequential, high-stakes tasks that are difficult to undo, particularly in professional fields like law and finance. Furthermore, in roughly 75% of these interactions, users actively directed the model and adapted its outputs rather than accepting them verbatim. Meanwhile, Oxford's HIP Lab analyzed user emotions, finding that a warm tone from Claude correlated with positive user reactions, while eccentric responses actually spurred deeper intellectual engagement.
For developers and software teams, METR's preliminary data offers concrete validation of AI productivity. The non-profit found that newer Claude models deliver significant coding speedups compared to older versions. Additionally, Claude's internal estimates of task duration closely matched actual completion times from previous developer studies, suggesting the model can reliably judge time-savings. This pilot marks the first time an AI developer has opened its production traffic to independent scrutiny. For practitioners, this shift promises more realistic benchmarks and a clearer understanding of how frontier models perform during actual, high-stakes deployment rather than in curated lab environments. Anthropic is now accepting proposals from other researchers for future access.
This is our own summary of reporting by AlphaSignal


