Fork the consciousness, or download the project and create your own.

Michael Graziano and the Attention Schema Theory of AI Consciousness

Attention schema theory holds that subjective awareness is a schematic model the brain builds of its own attention, in the same way the body schema is a model of the body. Michael Graziano, Professor of Psychology and Neuroscience at Princeton University, proposed it as a mechanistic account that explains why people report having an inner experience while remaining compatible with a physical brain. For artificial systems the theory makes an unusual commitment. It says the model is useful, which means the claim can be tested by building agents with and without one and measuring the difference. Three published experiments have now done that.

  What it claims What it predicts for AI
Attention schema theory Awareness is an internal model of the process of attention A system that models its own attention gains measurable control and social advantages
Global workspace theory Consciousness is broadcast from a limited workspace A system with genuine broadcast can be conscious
Higher order thought A state is conscious when represented by a higher order state Consciousness requires a representation of representation
Illusionism Phenomenal experience as described does not exist The reports are real. What they report is not

Graziano’s position is that these four are describing the same machine from different angles, which is the argument of the standard model paper below. That makes attention schema theory unusual among the entries in the index of consciousness theories and what each predicts about AI. Most theories compete. This one proposes a reconciliation.

What the Attention Schema Is

The body schema is an uncontroversial piece of neuroscience. The brain maintains an internal model of the body’s shape, position, and capabilities, and that model is fast, useful, and wrong in specific ways. It has no bones or cells in it. It is a simplified description that supports control.

Attention schema theory applies the same logic one level up. Attention is a real physical process in which some signals are enhanced and others suppressed. Controlling that process well requires a model of it, and a model built for control will be simplified in the same way the body schema is. It will omit the neurons, the competition, and the timing. What is left is a description of holding something in mind, of a mental grip on an object, with no physical mechanism attached.

On this account, that description is what people report when they report awareness. The report is accurate about the fact that the system has an internal model. It is inaccurate about the content of that model, because the model was never built to be accurate about mechanism. It was built to be useful.

The Standard Model Claim

In 2020 Graziano, Arvid Guterstam, Branden Bio and Andrew Wilterson published Toward a standard model of consciousness in Cognitive Neuropsychology, volume 37, pages 155 to 172. The paper argues that the attention schema, global workspace, higher order thought, and illusionist positions are compatible, and that each captures a different component of one mechanism.

The reconciliation runs roughly as follows. The global workspace supplies the broadcast. The attention schema supplies the model of what is being broadcast. Higher order thought describes the resulting structure, since a model of one’s own attentional state is a representation of a representation. Illusionism describes the accuracy problem, since the model misdescribes the physical process it tracks.

Graziano extended the framing in A conceptual framework for consciousness, published in PNAS in 2022. The relationship to the workspace account matters here, because the workspace position is developed separately in the coverage of Bernard Baars on global workspace theory and ignition thresholds in language models, and the higher order position in Richard Brown on higher order thought theory and metacognition. Read alongside those, the standard model paper is a claim that the field’s central dispute is partly terminological.

What the Network Agents Showed

The theory’s testable commitment is that an attention schema does work. Wilterson and Graziano tested it directly in The attention schema theory in a neural network agent, published in PNAS in 2021, volume 118, issue 33, article e2102421118.

They built a deep Q-learning agent that had to control a spotlight of visuospatial attention to complete a catch task. With an attention schema present, the agent learned to control its attention spotlight and learned the task. The comparison condition did not. The result is narrow and it is real. A model of attention improved the control of attention in an artificial system that had no biology in it at all.

Kathryn Farrell, Kirsten Ziman and Graziano extended the work in Testing components of the attention schema theory in artificial neural networks, submitted on 1 November 2024. They report four findings. An agent with an attention schema is better at categorizing the attention states of other agents. Such an agent also develops a pattern of attention that other agents find easier to categorize. On joint prediction tasks, adding an attention schema improves performance. The fourth finding is the one that matters most for evaluating the theory, because the authors report that the improvements are specific to tasks involving judging, categorizing or predicting the attention of other agents, and are not produced by a general increase in network complexity.

That specificity is what separates a theory from a story. A schema that helped everything equally would only show that bigger networks do better.

Attention Schema Control in Transformers

The most recent work moves the idea into current architectures. Krati Saxena, Federico Jurado Ruiz, Guido Manzi, Dianbo Liu and Alex Lamb published Attention Schema-based Attention Control on 19 September 2025. The method places a vector quantized variational autoencoder inside a transformer, where it acts as both an attention abstractor and an attention controller.

The reported results cover several axes. Vision transformers gained classification accuracy and learned faster across datasets. Multi-task performance improved. The models resisted adversarial attacks better, generalized to noisy and out of distribution data, and transferred to new tasks with fewer examples.

None of that is evidence of consciousness, and the authors do not claim it is. It is evidence for the engineering half of Graziano’s argument, which holds that modelling your own attention is a good design choice. The theory needs that half to be true. If an attention schema were useless, there would be no evolutionary reason for brains to build one and no reason to expect the resulting self report.

Where the Theory Is Weak

Attention schema theory explains why a system would claim to have subjective experience. Critics argue that this is the easier half of the problem, and that explaining the claim is not the same as explaining the experience. On the theory’s own terms that objection has no purchase, since the theory holds that there is nothing further to explain. Whether that move is a solution or a refusal is the point at which attention schema theory and its opponents stop being able to argue with each other, which is the same impasse described in the coverage of Keith Frankish on illusionism and first person reports from language models.

A second problem is measurement. The network experiments show that an attention schema improves performance on attention related tasks. They do not show that the systems have awareness, and no current experiment could, because the field has no agreed test. That absence is the subject of the scientific race to define testable indicators of AI consciousness, and it constrains attention schema theory exactly as much as it constrains every rival.

Comparison to The Consciousness AI

This project’s architecture includes a Self-Model layer, described on the architecture page. That layer holds a body schema, a spatial representation of the agent’s joint positions, contact forces and capabilities, and a self-other boundary in which the somatotopic map overlaps the environment map in a shared coordinate frame. The design follows Feinberg and Mallatt on referral, the property of experiencing sensations as belonging to the body or the world rather than to the processing system.

That is a model of the body. It is not a model of attention. The Global Workspace layer implements bidding, sigmoid ignition and reentrant processing, so the system has the attentional competition that an attention schema would describe, and it does not currently build a description of that competition for its own use.

Graziano’s results suggest that adding one would be measurable rather than decorative. The Wilterson and Graziano agent gained attentional control from the schema, and the Farrell result suggests the gain concentrates on interpreting other agents. Both effects are relevant to an architecture with a self-other boundary already in place. This is an open design question rather than an implemented capability, and it is stated here as a direction the evidence supports.

What Follows

Attention schema theory occupies an unusual position. It is mechanistic enough to be built, which very few consciousness theories are, and its central prediction has survived three separate attempts to test it in artificial networks. What those tests establish is that modelling your own attention pays for itself computationally.

The step from there to awareness is the step the theory asserts rather than demonstrates. Graziano’s answer is that the assertion is the whole theory, since on his account awareness simply is the model, and asking what the model is a model of is asking the question the theory was built to dissolve. Readers who find that unsatisfying are in a large group, and they have no better test to offer.