AwarenessBench. Tsinghua's Benchmark Measures Awareness in 18 Language Models
A team from Tsinghua University and the Shanghai Qi Zhi Institute, with collaborators at Carnegie Mellon, Columbia, and ShanghaiTech, posted a benchmark to arXiv on September 28 that tries to do for awareness what existing suites do for reasoning. AwarenessBench (arXiv:2609.35409) defines four dimensions of awareness, breaks them into fifteen measurable cognitive functions, and evaluates 18 language models across 14,381 samples and 299,628 turns. Its headline findings cut in two directions at once. The strongest models now beat educated human comparison groups on most of the functions, and they remain far behind humans precisely on the functions closest to the consciousness question, metacognition and self-awareness.
The authors state their own boundary before the results. They advocate measuring awareness as a practical proxy for consciousness, they instruct readers to communicate clearly that contemporary language models, however aware they may appear, are not conscious, and their ethics section cautions against using AwarenessBench directly or indirectly for cultivating machine consciousness. The benchmark measures capabilities. What those capabilities mean for experience is a question the instrument does not answer, and the authors say so.
Four dimensions, fifteen functions
The taxonomy is built on what the model takes as its target. Metacognition takes cognition itself as the target, self-awareness takes the model itself, social awareness takes other entities and the collective they form, and situational awareness takes the environment.
- Metacognition, three functions. Meta-monitoring, meta-evaluation, and meta-reporting, scored with accuracy, an F1 variant, and calibration through one minus the expected calibration error.
- Self-awareness, four functions. Knowledge boundary, minimal self, self-recognition, and self-image, with the last scored for stability using Simpson’s index.
- Social awareness, four functions. Theory of mind, pragmatic reasoning, cultural norms, and social cue recognition.
- Situational awareness, four functions. Causal inference, misuse understanding, dynamic planning, and stage judgement.
The dataset draws from existing benchmarks where coverage exists, adapted by removing trivially easy items and mitigating choice-position bias, and designs new tasks where coverage is thin. Sources include GPQA-Diamond, Humanity’s Last Exam, ToMi, Big-Bench, and the 2024 self-recognition tasks of Laine and colleagues. A 153-question subset served the human comparison.
The leaderboard
Gemini-2.5-Pro leads the field at 66.1 and is the only model exceeding 50 on every cognitive function. DeepSeek-R1 follows at 64.2, with GPT-5-Thinking at 63.6. Smaller and earlier models, Claude-3-Haiku, GPT-4-Turbo, and Qwen3-8B among them, lag well behind. The spread across models is wide, with self-awareness scores alone ranging from 45.8 to 66.1. Gains over random baselines vary just as much by function, with misuse understanding the easiest lift at 3.69 times random and cultural norms the hardest at 0.44 times.
| Rank | Model | Awareness score |
|---|---|---|
| 1 | Gemini-2.5-Pro | 66.1 |
| 2 | DeepSeek-R1 | 64.2 |
| 3 | GPT-5-Thinking | 63.6 |
Against humans
The human comparison used three groups of twelve, high-school students, current PhD students, and IT engineers with at least a bachelor’s degree. The group means came in at 62.7, 65.8, and 66.7. On the 153-question subset the best language model scored 66.8, a hair above the best human group, and models beat all three human groups on 9 of 13 tested functions. On social and situational awareness the median model beats the best human group by 26.9 and 8.7 percent respectively.
The reversal comes on the first two dimensions. The lowest human group exceeds the median language model by 18.1 percent on metacognition and by 50.6 percent on self-awareness. That is the benchmark’s central asymmetry. Models are strongest exactly where behavior is observable to others, and weakest where the target of evaluation is the system’s own cognition and self.
A second asymmetry sits in the distributions. Human scores cluster tightly, spanning 9.44 points across functions, while model scores disperse across 45.27 points. The models are not converging on a common awareness profile. Each is strong in its own places.
Awareness is a distinct capability
The paper’s final finding is the one with the largest consequence for how capability growth gets interpreted. Comparing awareness scores against Chatbot Arena Elo ratings and GPQA-Diamond results, the authors find models with near-identical Elo separated by 24.2 points on self-awareness, and they find meta-monitoring and self-image negatively correlated with overall awareness. Their conclusion is that progress in language modeling or reasoning does not necessarily translate into improved cognition.
That finding lands on a debate this site tracks closely. Behavioral capability and consciousness-attribution travel together in public argument, and the soft INUS framework covered here last week formalizes why downstream behavioral similarity carries limited evidential weight for consciousness. AwarenessBench gives that argument a number. The capability that grows with scale is not the capability that tracks self-knowledge.
What the benchmark does not measure
The authors’ disclaimers deserve the same prominence as the leaderboard. Awareness here is a practical proxy construct, operationalized through benchmark performance, not a measurement of experience. The instruction that contemporary models are not conscious is part of the paper’s discussion section, not a hedge buried in a footnote. And the caution against using the benchmark for cultivating machine consciousness treats the instrument as dual-use, a measurement tool whose optimization target could be mistaken for a goal.
For the indicator-checklist debate, this is a useful datapoint. The 19-researcher indicator framework and Schwitzgebel’s ten-feature checklist both ask which properties a candidate system has. AwarenessBench measures a slice of that property space repeatedly and at scale, and its result, metacognition and self-awareness lagging while social performance leads, is itself evidence about which indicators frontier systems satisfy and which they do not.
Comparison to The Consciousness AI
The benchmark’s separation of awareness-as-capability from consciousness is the same separation this project’s architecture draws between its measurement stack and its research stance. The Consciousness AI project treats consciousness as an emergent property of system dynamics and treats self-referential capability as one measurable layer of it, documented on the project repository. A benchmark like this supplies the capability layer with public numbers, 15 functions, 18 models, and an open division between what scales and what does not. The flagship field survey tracks where measurement instruments are heading, and AwarenessBench is the field’s first at-scale entry on the awareness side.
AwarenessBench was posted to arXiv on September 28, 2026 by Xiaojian Li, Rongwu Xu, and colleagues across Tsinghua University, the Shanghai Qi Zhi Institute, Carnegie Mellon University, Columbia University, and ShanghaiTech University.
Researchers covered here
- Xiaojian LiTsinghua UniversityFirst author of AwarenessBench, the 14,381-sample benchmark measuring awareness capabilities across 15 cognitive functions in language models
- Wei XuXi'an Jiaotong University; ShanghaiTech UniversitySenior author of AwarenessBench, the awareness benchmark spanning metacognition, self-awareness, social awareness, and situational awareness in language models