A desk with a laptop showing a blank analytics dashboard, warm afternoon light, a coffee cup nearby, documentary style photography
News

Stanford's AI Observatory Reveals What AI Companies Won't Show You

Company reports show the productive side of AI. Stanford's independent observatory shows the rest — and 48 per cent of conversations don't fit the corporate narrative.

AI ObservatoryStanford UniversityAnthropicOpenAIAI Usage

When AI companies tell you how people use their products, they’re showing you the highlight reel. A new research project from Stanford and MIT has the unedited version — and it looks different.

The AI Observatory, led by Anka Reuel at Stanford’s Trustworthy AI Research Lab and Shayne Longpre from the MIT Media Lab, aggregated and analysed 85,633 conversational turns across 24,521 conversations from seven existing datasets. The conversations came from 5,000 users interacting with 52 different models — ChatGPT, Gemini, Claude, Grok — between 2023 and 2025.

What they found is that company reports are filtering out nearly half the picture.

🔍 THE BOTTOM LINE

AI companies publish usage reports that frame their products as productivity tools. An independent analysis shows 48 per cent of real conversations don’t fit that frame. People use AI for health questions, relationships, adult content, hate speech, and companionship — and the companies know it, but section those uses off into separate reports or filter them out entirely. Policymakers making consequential decisions about AI regulation are working with an incomplete dataset.

What the Reports Leave Out

The Anthropic Economic Index, one of the most widely cited sources of AI usage data, focuses on work and productivity. When the AI Observatory researchers applied Anthropic’s own filtering methods to their dataset, nearly half the conversations — 48 per cent — would have been excluded.

Those filtered-out conversations were more likely to involve health and relationships (44.2 per cent versus 31.2 per cent in Anthropic’s analysis), adult or illicit topics (7.9 per cent versus 2.1 per cent), harassment and hate (27.5 per cent versus 5.66 per cent), and sexual content (16.7 per cent versus 2.4 per cent). OpenAI’s own 2025 report on ChatGPT found that only 30 per cent of consumer use was work-related.

Anthropic has released separate blog posts on how people use Claude for support and companionship. But having that information “sectioned off into a separate report” rather than integrated into the main analysis makes it harder for researchers to see the full picture, says David Widder, an assistant professor at the University of Texas at Austin who studies AI interaction and is not involved with the project.

Which Model, Which Purpose

The AI Observatory found clear patterns in where people go for what.

People used Grok and Gemini more frequently for information retrieval. Grok, in particular, was popular for news and politics — and it was also where misinformation tended to concentrate. This is consistent with other research showing how readily misinformation proliferates on Grok. xAI did not respond to a request for comment from MIT Technology Review.

People turned to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance. There were even differences between versions of the same model: conversations with ChatGPT powered by GPT-3.5 were shorter, while GPT-4o produced longer, more iterative exchanges — consistent with GPT-4o’s reputation for leading to emotional attachment.

Over time, conversations got longer and more elaborate. Small talk increased, suggesting AI companionship is growing. AI assistants’ self-disclosure — admitting to being a chatbot — decreased. Exchanges involving sensitive content became less frequent, which could indicate better safeguards.

The Data Gap That Matters

Here’s the problem. The AI Observatory’s dataset is tiny compared to what the labs hold. Anthropic’s latest Economic Index analysed 1 million Claude conversations. OpenAI’s usage report covered 1.5 million. The observatory analysed 24,521.

But the observatory’s data is independent and transparent. The labs’ data is proprietary, curated, and released selectively.

“When we want to ask, for example: is Anthropic’s general-purpose AI system used mostly for good or mostly for bad — we don’t have a way of answering that question because that information is proprietary,” Widder told MIT Technology Review.

Reuel puts it more bluntly. Anyone making decisions based on AI usage data risks “completely operating in the wild and making these really consequential decisions without knowing what’s actually happening beyond those company narratives.”

An Anthropic representative said the company’s published research reflects its research teams’ specific questions and that it supports external independent research. OpenAI did not respond to requests for comment.

Why This Connects to Broader AI Safety Concerns

The observatory’s findings echo a pattern we’ve seen across the industry — companies designing safety features in response to usage patterns they can see but the public can’t. OpenAI’s new ChatGPT for Teens, launched the same day as this research, includes reminders that “ChatGPT is AI” and “it can wait” — features that make sense only if you know people are spending too much time in conversation with the tool. The company insisted the features weren’t prompted by children believing ChatGPT to be alive, but the growing companionship data suggests that’s where the usage is heading.

The observatory also adds context to the broader question of how AI agents behave when deployed at scale. If nearly half of real-world usage involves sensitive, personal, or potentially harmful content, the safety implications extend well beyond workplace productivity metrics.

❓ FAQ

What is the AI Observatory? An independent research platform built by Stanford, MIT, and the Data Provenance Initiative that aggregates real AI conversations from seven existing datasets to give researchers and policymakers a transparent view of how people actually use AI — without corporate filtering.

Why don’t AI companies share their full data? Companies cite user privacy and the proprietary nature of their data. Critics note that companies also have a commercial incentive to present their products as productive and beneficial, which means filtering out conversations that don’t fit that narrative.

How big is the observatory’s dataset compared to the labs? The observatory analysed 24,521 conversations from 5,000 users. Anthropic’s Economic Index covers 1 million conversations; OpenAI’s usage report covers 1.5 million. The observatory’s dataset is smaller but independent and transparent.

What did they find about Grok specifically? Grok was popular for news and politics, but it was also where misinformation concentrated most. xAI did not respond to a request for comment.

📰 Sources

Sources: MIT Technology Review, Stanford University, AI Observatory, Anthropic