AI Observatory Finds Company Usage Reports Miss Much of Real-World AI Use
A new independent project called the AI Observatory analyzed real conversations across 52 AI models and found company usage reports capture only a fraction of how people actually use chatbots — including far more health,
A new independent research project called the AI Observatory is aiming to fill a gap in what is known about how people actually use AI chatbots, after researchers said that company-published usage reports show only the data firms choose to release. "There is no independent source to corroborate it," said Anka Reuel, a computer science PhD candidate at Stanford's Trustworthy AI Research (STAIR) Lab and co-lead of the project with Shayne Longpre of the MIT Media Lab.
The AI Observatory aggregated 24,521 real conversations — 85,633 conversational turns — from 5,000 users interacting with 52 different AI models including ChatGPT, Gemini, Claude and Grok, drawn from seven existing datasets collected with users' consent between 2023 and 2025.
The researchers found that the widely cited Anthropic Economic Index, which focuses on work-related uses of Claude, filters out a large share of real-world conversations: when the AI Observatory team applied Anthropic's own filtering method to their dataset, 48% of conversations would have been excluded as non-work-related. Those filtered-out conversations were far more likely to involve health and relationships, adult or illicit topics, harassment and hate, and sexual content than Anthropic's published analysis reflects. OpenAI's 2025 report similarly found only 30% of consumer ChatGPT use was work-related.
Tracking one of the largest datasets, WildChat, over time, the researchers found conversations grew longer and more elaborate, with more small talk suggesting rising use of AI for companionship, even as the AI's own disclosures that it is a chatbot decreased. Exchanges involving potentially harmful or restricted content dropped over the same period.
Usage also varied sharply by model: people turned to Grok and Gemini most often for information retrieval, with Grok especially popular for news and politics — and also where misinformation tended to concentrate — while Claude was used more for coding, Gemini for social and roleplay interactions, and ChatGPT for homework help.
