Assessing Mentalization in Humans and Large Language Models
Language models produce behavior consistent with humans on theory-of-mind tasks. Whether they can use mentalization to guide adaptive behavior is a separate question, and a new preprint is designed to test exactly that. The authors run two economic games with cognitive computational modeling to uncover the latent strategies behind model behavior, and benchmark the results against a human sample.
The paper by Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, and Lei Zhang was posted to arXiv as 2608.26291. The team measured 2,099 individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5, and Gemini 2.0 Flash, against opponents of varying sophistication, and compared them with 251 human participants.
Two games, one latent profile
The design avoids the usual weakness of theory-of-mind benchmarking for models, which is that text responses can be rehearsed. Economic games force choices under strategic uncertainty, and computational modeling of those choices reveals the latent strategy, whether the agent tracks the opponent’s beliefs and intentions or presses a simpler rule.
Across both games, the models showed clear behavioral and computational signatures of mentalizing, and those signatures differed markedly by model provider and size. The result is not a single “models can or cannot mentalize” verdict. It is a graded map of a capability that some model families instantiate more consistently than others.
| Family | Signature | Notable result |
|---|---|---|
| DeepSeek, Gemini 2.0 Flash | Mentalizing present, weaker | Provider-dependent profile |
| GPT-4.1 | Mentalizing present, middle | Provider-dependent profile |
| GPT-5 | Strongest mentalizing | Adapted recursion depth, beat humans |
Strategic prompting and recursive depth
The authors examined whether a prompting strategy designed to elicit strategic reasoning improved performance. Strategic prompting generally did, by inducing more sophisticated reasoning, but the benefit differed across the two tasks. This asymmetry is itself informative: a model can reason more deeply when prompted without that meaning its default behavior mentalizes.
The standout result is GPT-5. It flexibly adapted its recursive depth of reasoning to increasingly sophisticated opponents and outperformed the human participants. The authors phrase it as demonstrating different capacities for mentalization across LLMs, and they position cognitive computational modeling as a formal method for assessing comparative intelligence across humans and machines.
Why this matters for consciousness research
The site has consistently treated reported “theory of mind in LLMs” claims with care. The Butlin indicator rubric includes social and metacognitive structure among its markers, and the site’s philosophical stance on what a conscious agent is separates capability from experience. This paper is close to that distinction in method. It measures a latent capability structure, not experience, and it does so on a shared formal footing for humans and machines.
The paper also strengthens the case that a single benchmark number for “mentalization” is meaningless. The provider-by-provider and size-by-size variance is the finding, not the average. That variance is exactly what the site argues should be reported honestly in the field, and it is consistent with the state-of-field consensus article, which notes that no current system meets established behavioral indicators in full.
Limits
The preprint has not passed peer review. The games, opponent models, and prompt conditions bound the result. The claim is about latent strategy structure in a defined setting, not a general statement about minds. What is established is a method, a shared instrument for humans and machines, and a displacement of the binary question with a measured, family-dependent profile.