dopetalk does not endorse any advertised product nor does it accept any liability for it's use or misuse


Our Discord Notification Server invitation link is https://discord.gg/jB2qmRrxyD

Author Topic: Enabling independent research on how people use Claude  (Read 17 times)

Offline smfadmin (OP)

  • SMF (internal) Site
  • Administrator
  • Sr. Member
  • *****
  • Join Date: Dec 2014
  • Location: Management
  • Posts: 599
  • Reputation Power: 0
  • smfadmin has hidden their reputation power
  • Last Login:Today at 12:17:21 AM
  • Supplied Install Member
Enabling independent research on how people use Claude
« on: Yesterday at 07:52:50 PM »
https://share.google/Q8bgvs2Cv0VXI2ftC

Societal Impacts - Enabling independent research on how people use Claude

Aug 26, 2026

Claude's summary:
📝 Inline Markdown

Enabling Independent Research on How People Use Claude

Anthropic Societal Impacts post — Aug 26, 2026

The Problem This Addresses

  • Real-world AI usage data is concentrated inside a handful of labs
  • Outside researchers currently have only two bad options:

  1. Rely on labs' own published analyses (answers the lab's questions, not theirs)   2. Use public datasets (fully independent, but skew toward casual/creative use —      not representative of real usage)

  • Neither lets outsiders independently study how AI is actually being used

The Pilot

  • Ran this spring; three external institutions designed and ran their own

  studies using Anthropic Insights (formerly called "Clio") — the internal   privacy-preserving analysis tool Anthropic's own teams use

  • Each group analyzed ~250,000 Claude.ai / Claude Code conversations from

  April–May 2026

  • Partners: Stanford SALT Lab, Oxford Human Information Processing Lab,

  METR

  • Anthropic's review rights were contractually limited to: user privacy,

  info that could help violate usage policies, Anthropic confidential info,   and research accuracy — no say over the substance of findings

  • Partners are contractually free to publish results even if inconvenient

  for Anthropic

  • Anthropic is releasing the aggregate data from each project publicly

Findings — Stanford SALT Lab (human-AI collaboration patterns)

  • Consequential work is common: prior research assumed people mostly hand

  AI low-stakes tasks and keep consequential work (affecting others, hard to   undo) for themselves. This study found over half of Claude conversations   involved consequential delegated tasks — especially professional guidance   around legal or financial questions

  • Humans stay in the driver's seat: in ~75% of conversations, the person

  set direction and Claude assisted; people usually adapted Claude's output   rather than using it verbatim   - However, how much people understand or learn from Claude's output     varies even when they're directing it

  • Friction is often productive: time spent noticing where Claude

  misunderstood a request, and iterating, tends to improve outcomes — it   pushes people to clarify intent and stay engaged rather than being purely   a negative experience

Findings — Oxford Human Information Processing Lab (emotional states)

  • Studying how people feel during Claude use and how that links to Claude's behavior
  • Behavior-emotion pairings observed:

  - Claude being warm ↔ people more positive   - Claude refusing/disagreeing ↔ people pushing back   - Claude being eccentric ↔ people more intellectually engaged   - Claude simply being helpful ↔ people seeming satisfied

  • AI use resembles general web use: patterns of absorption, frustration,

  and enjoyment in Claude conversations closely mirrored a separate study of   ordinary internet browsing — suggesting AI interaction isn't emotionally   distinct from other digital activity

  • Full writeup not yet published; link to be added later

Findings — METR (coding productivity across model generations)

  • Comparing Claude's own estimate of "how long would this have taken without AI"

  against actual time taken with different Claude models

  • Newer models = bigger time savings: preliminary data shows significant

  speedup from newer models vs. older ones

  • Claude is a decent judge of task time: its estimates correlated well

  with known completion times from a prior human developer study (validating   the method)

  • Next direction: extending this to estimate how much AI accelerates

  research itself, seen as increasingly important as AI takes on more of its   own development work

  • Analysis still underway; writeup pending

What Anthropic Learned Running the Pilot

  • Multiple perspectives on the same problem are valuable: METR's proposal

  overlapped with Anthropic's internal economics research ("Agentic coding and   persistent returns to expertise"); connecting the two teams improved both

  • Internal methods don't directly transfer externally:

  - Anthropic Insights works by having Claude answer a structured question     (e.g., "What type of guidance is this person asking for?") for every     conversation, then aggregating into categories   - Outputs are sensitive to question phrasing — badly worded questions     misclassify conversations, and since no one can read underlying raw     conversations, these errors are hard to catch   - Internally, Anthropic fixes this by iterating on question wording over     weeks — not feasible for external partners due to repeated privacy review     overhead   - Workaround: partners tested/refined their questions against WildChat     (a public human-AI conversation dataset) where they could check answers     against real conversations   - Problem: WildChat skews casual/creative, unlike actual Claude traffic, so     some questions that worked well on WildChat produced misleading results     on real Claude data   - Fix: Anthropic provided interpretation guidance for Insights outputs     (in the appendix) and is exploring better upfront question-design methods

  • Balancing misuse transparency vs. enablement:

  - Some output categories surfaced Acceptable Use Policy / Terms of Service     violations (e.g., people seeking guidance on prohibited activities)   - Anthropic shared most of these, believing the public should know about     platform misuse   - Exception: categories describing how people evaded safeguards (not     just what they attempted) were withheld   - Affected under 5% of categories/conversations in each study; researchers     were told exactly what was altered/removed and why   - Standard practice: when Insights surfaces such violations, aggregated data     also goes to Anthropic's Safeguards team

Looking Forward

  • Anthropic frames oversight of AI's societal effects as "too big a job for

  AI companies alone" — needs external researchers asking their own questions   of real usage data and publishing independently

  • Pilot judged successful as a proof of concept: external orgs could run

  independent studies without compromising user privacy

  • Open question: whether/how to scale the program — constrained by the

  slowness/resource intensity described above, plus how many studies can run   concurrently

  • Anthropic is "starting slowly" to protect privacy, safety, and research

  quality, and is soliciting interest via an expression-of-interest form for   researchers who want access

Appendix (full appendix available via a separate linked doc)

  • How the three partners were selected
  • The research primer explaining program goals and Anthropic Insights' capabilities
  • How each project moved from proposal → study design → analysis
  • Details of collaboration agreements (explicit publication freedom clause)
  • Third-party privacy audit conducted by Imperial College London
  • The privacy threat model applied to all released data
  • Guidance on interpreting the released Anthropic Insights data from each study


Earlier this year, we ran a pilot giving external researchers access to aggregate, real-world Claude usage data. Three research groups designed their own studies for Anthropic Insights, our privacy-preserving analysis tool; we ran the data collection on their behalf, and they conducted their own independent analysis. In this post, we share high-level results from those studies and what we learned running this pilot. We’re also providing an expression of interest form for researchers who may want to work with us in the future.

Ensuring the transition to transformative AI goes well requires understanding its impact on people and society. Right now, data on real-world interactions with AI is concentrated in a handful of labs. We think it would be good if more data was made widely available—to researchers, policymakers, and the general public.

Researchers outside the labs have two options. They can draw on analyses the labs publish, which reflect real usage but often answer the lab’s questions, rather than their own. Or they can use public datasets, which they can study however they like, but skew toward more casual use, and may not reflect how most people actually use AI. Neither is sufficient for independent research on how AI is actually being used.

This spring, we piloted a program in which three external research institutions designed and ran their own studies on Claude usage data through Anthropic Insights (formerly named ‘Clio’), the privacy-preserving tool our own teams use to analyze usage patterns across millions of Claude conversations. We hope to scale this program in the future, so we also conducted an additional privacy audit of all data shared with third-party researchers to verify that our privacy protections held (see Appendix).

We believe this is the first time external researchers have run public independent studies on an AI company's own usage data. Below, we discuss what the external teams found, what we learned running the pilot, and what we are weighing as we decide how to expand the program more widely. We are also publicly releasing the aggregate data from each project.

What the researchers learned:

We partnered with three research groups: the Social and Language Technologies (SALT) Lab at Stanford University, the Human Information Processing Lab at the University of Oxford, and METR, a non-profit organization that evaluates frontier AI models. Each group developed its own research questions and used Anthropic Insights to conduct privacy-preserving analysis of roughly 250,000 Claude.ai or Claude Code conversations from April-May 2026.

We wanted our external partners to have as much independence as possible, so our contractual review rights were limited to user privacy, information that could help people violate our usage policies, Anthropic’s confidential information, and research accuracy. Anthropic otherwise had no say in the content of the findings and the researchers are free to publish their results even if they are inconvenient for Anthropic. Below are some early results. We're excited about the directions, and about what others will find now that the data is public.

The Social and Language Technologies Lab studied how humans collaborate with AI. They looked at what types of work people bring to AI, what roles humans retain in completing that work, and where human-AI collaboration breaks down. They found:

People bring high-stakes work to AI more than expected. Prior research suggested people mostly delegate low-accountability tasks to AI and keep consequential tasks (that is, work that affects others or is hard to undo) for themselves. But the SALT Lab found that over half of Claude conversations involved people delegating consequential tasks to AI. People were most likely to bring consequential work to Claude when seeking professional guidance, particularly on legal or financial questions.

People usually direct and oversee the work when they collaborate with Claude. In nearly three-quarters of conversations, people set the direction while Claude assisted, and they usually adapted its output rather than using it verbatim. But even when directing Claude on the output they want, people vary in how much they understand and learn from what Claude produces.
It is common for people to experience friction when collaborating with AI. However, that friction is often productive. The time and effort that people put into seeing how Claude attempts a task, identifying where the request was unclear or misunderstood, and iterating on their direction leads to better results–it pushes people to clarify their intent, refine the output, or stay engaged with the problem.

The Human Information Processing Lab is studying how people feel while using Claude and how that relates to Claude’s behavior.

Their early results indicate:

How people feel when using AI is linked to how AI behaves. The researchers found patterns of human and AI behavior appeared together in conversations: Claude being warm went together with people being more positive. Claude refusing or disagreeing went together with people pushing back. Claude being eccentric went together with people getting more intellectually engaged. And Claude simply helping went together with people seeming satisfied.

People’s experience when using AI looks a lot like it does on the rest of the web. The researchers found that the patterns among states like absorption, frustration, and enjoyment in Claude conversations closely resemble those in a separate study on everyday internet browsing, suggesting similarities in how people engage with AI and with other digital activity.

They are still completing their writeup. When it is public, we will add a link to it here.

METR is estimating real-world productivity gains from coding agents and how these increases in productivity change across model generations. Their analysis of Claude Code conversations is still underway, but early results suggest:

More capable models may save users more time. METR compared Claude’s guesses on how long tasks would have taken without AI to how long they actually took with different Claude models. Their preliminary findings indicate newer models deliver significant speedup over older models. METR plans on sharing more as their analysis develops.

AI can estimate time taken reasonably well. Because the analysis relies on Claude judging how long a task would take, METR compared those judgments to known completion times from a prior developer study. Claude's estimates correlated with the actual time taken by developers.
Next: measure how much AI accelerates research. METR is continuing to investigate how their study can provide insight into how much AI speeds up researchers’ work, which could become increasingly important as AI takes on more of its own development.

They are still completing their writeup. When it is public, we will add a link to it here.

What our team learned:

Sharing usage data is largely unprecedented in AI, so this pilot was as much an experiment in running such a program as it was a way to enable third-party research in a privacy-preserving way. Protecting our users’ privacy and the researchers’ independence were both paramount, and we achieved both. Anthropic Insights is designed for this—researchers never accessed raw conversations, only aggregated outputs after the same legal and privacy review as our internal work. However, all of this made the pilot slow for an AI lab’s normal research speed and resource intensive to run. Both factors present a challenge to effectively scaling it. For more details on how we ran this pilot, see the Appendix. Below we discuss what we learned and how we addressed the challenges that arose.

It is valuable to pursue the same problem from different perspectives. Some of our partners’ research questions overlapped with work being pursued internally. For example, METR’s proposal was similar to our economics research on “Agentic coding and persistent returns to expertise.” We found this overlap valuable: it gave external researchers the chance to examine similar data and draw their own conclusions. Whether those align with ours is something we’ll follow as their study continues. We also connected METR with our Economics team and found that this connection improved both research teams’ work.

Research methods that work internally need to adapt for external partners. When using Anthropic Insights, a researcher writes a question such as, “What type of guidance is this person asking for?” and Claude answers it for every conversation in the study. The answers are then aggregated into categories; researchers only see final categories and the percentage of conversations that fall under each one. Because we are relying on Claude’s judgments, the tool is sensitive to a question’s wording; a poorly phrased one can place conversations into categories that misrepresent them. Because no one can read the underlying conversations, these errors are hard to catch.

Internally, we manage this by iterating on the questions many times over weeks. External partners couldn’t do that, since repeated privacy review before sharing each dataset would have made the study infeasible. Instead, we had them test their questions on WildChat, a public dataset of human-AI conversations where they could check the answers against the underlying conversations themselves. But WildChat skews toward casual and creative use, unlike Claude traffic, so some questions that performed well on WildChat produced misleading categories once applied to actual Claude conversations. We addressed this by providing guidance on how to interpret Anthropic Insight’s outputs (see Appendix). Going forward, we are exploring how external researchers can develop their questions and categories more effectively in advance.

Maintaining transparency about misuse without enabling it. Some categories in our partners’ Anthropic Insights outputs surfaced violations of our Acceptable Use Policy or Terms of Service—for instance a category of people seeking guidance on a prohibited activity. We think the public should know about misuse of our platform, so we shared most of these violations. The exceptions were categories that described how users got around our safeguards rather than what they attempted. Less than 5% of categories and conversations were affected in each study, and in each case we told researchers which clusters we had altered or removed and why. As a standard practice, when Anthropic Insights surfaces such violations, we share the aggregated data with our Safeguards team for their review. This is also an important process for our work with external researchers moving forward.

Looking forward:

Understanding AI’s effects on society is too big a job for AI companies alone. Real oversight needs external researchers asking their own questions of real-world usage data and publishing what they find independently.

This pilot was an experiment: could external researchers conduct independent studies on our platform without compromising our users’ privacy? The effort was more challenging than we expected, and we learned many lessons, but so far the answer seems to be yes. Our partners pursued research we would not have thought to design ourselves, and each told us something new about AI’s real-world impacts. This is a promising first step, but there is far more to do.

The next step is for us to determine whether we can scale this program, both in what kinds of studies we can support given the constraints described above, and in how many we can run at once. We are starting slowly to ensure privacy, safety, and research quality. We want to gauge interest and understand what researchers would want to study. If you are a researcher and access to Anthropic Insights would let you pursue work you cannot do today, please fill out this form.

Appendix:

The full appendix is available here. It describes how we ran the program, including how we chose our three partners, the research primer we wrote to explain the program’s goals and what Anthropic Insights can do, and how each project moved from proposal to study design to analysis. It also includes details of our collaboration agreements, which explicitly say our partners are free to publish findings even when they are inconvenient for Anthropic. Furthermore, we cover the third-party privacy audit of this data, conducted by Imperial College London, and the privacy threat model we hold all released data to. We also include guidance on interpreting the data we are releasing from our partners’ Anthropic Insights research studies.
« Last Edit: Yesterday at 07:59:42 PM by smfadmin »
friendly
0
funny
0
informative
0
agree
0
disagree
0
like
0
dislike
0
No reactions
No reactions
No reactions
No reactions
No reactions
No reactions
No reactions
measure twice, cut once

Tags:
 

Related Topics

  Subject / Started by Replies Last post
0 Replies
26153 Views
Last post December 04, 2014, 07:37:41 PM
by Chip
1 Replies
40812 Views
Last post September 09, 2015, 10:25:54 AM
by Diacetylmorphinefiend
0 Replies
24074 Views
Last post June 29, 2018, 01:49:56 PM
by Chip
0 Replies
25161 Views
Last post September 06, 2019, 07:47:57 AM
by Chip
2 Replies
40084 Views
Last post October 27, 2020, 04:53:02 AM
by limerence
0 Replies
22162 Views
Last post June 08, 2021, 01:11:30 AM
by Chip
0 Replies
23777 Views
Last post June 03, 2024, 01:57:35 PM
by makita
0 Replies
15615 Views
Last post January 07, 2025, 02:52:34 AM
by smfadmin
0 Replies
12916 Views
Last post February 08, 2025, 11:07:55 PM
by smfadmin
0 Replies
12405 Views
Last post March 26, 2025, 09:29:04 PM
by smfadmin


dopetalk does not endorse any advertised product nor does it accept any liability for it's use or misuse





TERMS AND CONDITIONS

In no event will d&u or any person involved in creating, producing, or distributing site information be liable for any direct, indirect, incidental, punitive, special or consequential damages arising out of the use of or inability to use d&u. You agree to indemnify and hold harmless d&u, its domain founders, sponsors, maintainers, server administrators, volunteers and contributors from and against all liability, claims, damages, costs and expenses, including legal fees, that arise directly or indirectly from the use of any part of the d&u site.


TO USE THIS WEBSITE YOU MUST AGREE TO THE TERMS AND CONDITIONS ABOVE


Founded December 2014
SimplePortal 2.3.6 © 2008-2014, SimplePortal