Anthropic Gives Outside Researchers Access to Real Claude Usage Data
Anthropic opened privacy-preserved Claude usage data to Stanford, Oxford, and METR researchers for the first time, covering ~250,000 conversations from April-May 2026.

For the first time, Anthropic has let external researchers study real Claude usage data. Three independent teams at Stanford University, the University of Oxford, and the nonprofit AI evaluator METR each designed studies using aggregated data from roughly 250,000 Claude.ai and Claude Code conversations recorded between April and May 2026. Researchers never saw the underlying text. Instead, Anthropic's privacy-preserving system called Anthropic Insights ran their queries and returned category-level results. One completed study has already surfaced findings worth paying attention to.
What happened
| Detail | Fact |
|---|---|
| Conversations analyzed | ~250,000 Claude.ai and Claude Code sessions |
| Date range | April to May 2026 |
| Research teams | Stanford University, University of Oxford, METR |
| Privacy system used | Anthropic Insights (aggregated outputs only) |
| Privacy audit | Independent audit by Imperial College London |
| Categories altered or removed | Fewer than 5% per study |
Anthropic describes this as the first time it has opened real-world Claude usage data to people outside the company. The setup was deliberately constrained: researchers submitted questions such as “what type of guidance is this person seeking?” and Anthropic’s Insights system evaluated conversations against those questions, then returned grouped percentages. No one on the research teams could read the original conversations.
Anthropic’s review rights covered four areas only: user privacy, potential policy violations, confidential company information, and research accuracy. Outside those limits, researchers were free to publish results that reflected poorly on Anthropic. An independent privacy audit was conducted by Imperial College London.
What the Stanford study found
The only completed study so far came from Stanford’s Social and Language Technologies Lab (SALT Lab). Its headline finding: more than half of the Claude conversations analyzed involved tasks the researchers classified as “consequential,” meaning the work either affects other people or is difficult to reverse. Professional guidance, especially legal and financial questions, came up most often in this category.
In nearly three-quarters of conversations, users set the direction and Claude assisted. Users also generally adapted Claude’s output rather than taking it word for word. But researchers found meaningful variation in how well users appeared to understand the work Claude was producing, even when they were the ones steering the conversation.
The study also pushed back on a common assumption about friction. Misunderstandings and back-and-forth were common, but the researchers found that working through them often helped users clarify their goals and stay engaged rather than simply accept a flawed answer.
What Oxford and METR are studying
Two studies are still underway. Oxford’s Human Information Processing Lab is looking at how users appear to feel while interacting with Claude and how those states connect to the model’s behavior. Early findings suggest patterns on both sides: warmer Claude responses appeared alongside more positive user behavior, while refusals were associated with users pushing back. More unusual Claude outputs appeared alongside greater intellectual engagement from users. Researchers noted that states like absorption, frustration, and enjoyment resembled patterns seen in studies of everyday web browsing.
METR is using Claude Code conversations to study productivity gains from coding assistants and whether those gains shift as models improve. Preliminary results suggest newer Claude models may save users more time than earlier versions, though the study is not yet complete. If you’re already using AI in your development workflow, the GitHub Copilot HydraFusion rollout is a useful parallel to watch alongside this research.
Why this research method has real limits
Giving external researchers access to real usage data solves one problem (public datasets skew toward casual or creative use and may not reflect everyday patterns) but creates others.
- The analysis still depends on Claude itself making classification judgments.
- Small changes in how a research question is phrased can shift the output categories.
- Researchers cannot inspect original conversations, so misleading classifications are hard to catch.
- Questions that worked reliably on WildChat, a public human-AI conversation dataset, performed less consistently on Claude conversations because the two datasets have different usage patterns.
Anthropic says it is exploring ways for future research teams to develop and validate their queries more effectively before studies begin. The aggregated datasets from all three pilot projects are being released publicly.
Why it matters
AI companies have historically kept usage data inside the building. That is partly commercial, partly legal, and partly technical. What gets published tends to be cherry-picked demos or safety red-teaming results. This pilot is a meaningful step toward independent verification of how these models are actually used, not just how they perform on benchmarks.
The Stanford finding about consequential tasks is worth sitting with. If more than half of real Claude conversations involve work that is hard to undo or that affects third parties, then the stakes of a confidently wrong AI answer are higher than they might seem during a casual demo. This is relevant context for any business considering AI integration for client-facing or regulated workflows.
It also matters for the broader question of AI accountability. Right now, Anthropic controls the Insights system, sets the review criteria, and decides what gets removed before researchers see results. Fewer than 5% of categories were altered in each study, and researchers were told what changed, but outside researchers still cannot independently verify those claims.
Our take
This is a genuinely useful pilot, and the Stanford results alone justify the effort. Knowing that most Claude users are running consequential tasks, not just casual queries, changes how you think about error rates and oversight.
That said, the methodology has a structural problem that Anthropic acknowledges: you are asking an AI model to classify AI conversations, and the researchers cannot audit the raw material. That is not a fatal flaw, but it means these findings should be treated as directional, not definitive. The WildChat validation step was smart, and the fact that it exposed reliability gaps before publication is a good sign for rigor.
For business owners thinking about how their teams are using AI tools right now, the Oxford emotional-state research is the one to watch longer term. If there is a meaningful link between how an AI responds and how engaged or frustrated users become, that has practical design implications for anyone building AI-assisted products or internal tools.
Anthropic describes this pilot as resource-intensive and slower than its internal research process. Broader access is not guaranteed at scale, but the company is collecting expressions of interest. If independent AI research is part of your work, this is worth following. And if you want a clearer picture of how AI tools are actually performing inside your own business, that starts with tracking your own data, which is something our team covers as part of our broader AI and automation services.
Frequently asked questions
What is Anthropic Insights?
Anthropic Insights is Anthropic's privacy-preserving analysis system. Researchers submit questions about Claude conversations, the system evaluates those conversations against the questions, and returns aggregated category breakdowns. Researchers never see the underlying raw conversations.
What did the Stanford Claude usage study find?
Stanford's SALT Lab found that more than half of Claude conversations involved consequential tasks, meaning work that affects other people or is difficult to reverse. Professional guidance such as legal and financial questions was particularly common. In nearly three-quarters of conversations, users directed the work while Claude assisted.
How many Claude conversations were analyzed in the Anthropic research pilot?
Roughly 250,000 Claude.ai and Claude Code conversations, recorded between April and May 2026, were used across the three research studies.
Can outside researchers now freely access Claude usage data?
Not yet at scale. Anthropic describes the pilot as resource-intensive and slower than its internal research process. The company is currently collecting expressions of interest from researchers who want to conduct studies requiring access to real AI usage data.


