
Additionally, exchanges that the researchers labeled as sensitive use—meaning ones with potentially harmful or restricted content, including sexual harassment and hate speech—dropped. That might suggest that platforms were generally deploying more effective safeguards.
The AI Observatory also found that AI use looked significantly different depending on the model. Depending on the tool, users ranged in topics, interaction styles, conversation structures, as well as both the likelihood and type of sensitive use cases.
For example, the researchers found that people used Grok and Gemini more frequently for information retrieval. Grok, in particular, was especially popular for information on news and politics, but it was also where misinformation tended to concentrate. (This is consistent with other research that has also shown how readily misinformation proliferates on Grok. xAI did not respond to a request for comment.)
Meanwhile, people were more likely to turn to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance.
There were even differences among different versions of the same model. Researchers found that people had shorter conversations with ChatGPT when it was powered by GPT-3.5, and longer and more iterative ones with GPT-4o—which makes sense given that that version became known for leading to emotional addiction.
Company reports, however, didn’t tend to capture these nuances across or even within their own models. “No single company report tells the whole story,” says Shayne Longpre, a recent PhD graduate from the MIT Media Lab who co-led the research with Reuel.
To create the AI Observatory, Reuel and researchers from MIT, Stanford, the Data Provenance Initiative, and other institutions, aggregated 24,521 conservations across 85,633 conversational turns (that is, the user prompt and corresponding AI response) from seven real-world datasets collected by previous research. These conversations came from 5,000 users interacting with 52 different models, including ChatGPT, Gemini, Claude, and Grok, between 2023 and 2025.
But these conversations are a drop in the proverbial bucket compared to the data that the big labs themselves have access to. The latest Anthropic Economic AI Index, for example, is based on analysis of 1 million Claude conversations; OpenAI’s report on how people are using ChatGPT analyzed 1.5 million conversations.