Fork the consciousness, or download the project and create your own.

Afshin Khadangi Causal Liability Theory and the AI Consciousness Fallacy

Afshin Khadangi published “We Built a Mirror and Mistook It for a Mind” on arXiv on September 6, 2026 (arXiv:2609.06715). The paper argues that the entire debate over machine consciousness starts from an unexamined assumption. When a language model produces sentences that look like reports of inner experience, debaters on both sides treat the model as the kind of entity to which consciousness could belong. Khadangi names that inference the AI Consciousness Fallacy. A generative system can return linguistic traces of human interiority in first-person form without there being any phenomenal bearer behind the words. His response is Causal Liability Theory, a framework that separates three things the debate routinely collapses, and then defines a criterion for when a physical process qualifies as a candidate bearer at all.

This is a contribution to consciousness measurement in the strict sense. The paper does not claim that models are conscious. It does not claim they are not. It separates attribution, individuation, and constitution into three independent questions, and it supplies an experimental protocol for the second one.

Three Things Debaters Keep Collapsing

Khadangi’s opening move is a three-way separation. Phenomenal consciousness, the existence of subjective experience. Introspective report, the linguistic output a system produces about itself. Human projective introspection, the reader’s own inner states reflected back into the text. A language model trained on human writing absorbs an enormous inventory of first-person idiom. When it outputs “I feel” or “I notice,” the output is a statistical trace of the human interiority in its training distribution. The first-person grammatical form carries no evidence that the system itself is the subject of any experience.

The fallacy lies in the reverse inference. Because the words look like reports, observers conclude there is a report-generating subject. Khadangi’s point is that the inference requires an additional premise, the existence of a bearer, and that premise is exactly what nobody has tested. His framework is built to make it testable.

The separation connects directly to the consciousness vector work by Kim, Street, Rocca, Waytz, Evans and Keeling. That study located an internal direction in language model activations encoding the model’s stance on its own phenomenal status, and showed the direction has downstream causal effects. A mechanistically real self-stance is not a bearer. Khadangi’s framework says the two must be distinguished before any measurement can be interpreted, and the CLT audit operationalizes the distinction.

Liability Closure

Causal Liability Theory comes in two parts. CLT-I proposes a criterion for individuating a candidate bearer. Khadangi calls the criterion liability closure. A physically continuing process becomes the non-delegable inheritor of constraints generated by its own endogenous discriminations. In plainer terms, a system qualifies as a bearer candidate when its own internal discriminations generate constraints that the same continuing process, and nothing outside it, must carry forward. The process cannot delegate its constraints elsewhere and remain the same bearer.

The criterion is deliberately substrate independent. Nothing in its statement mentions neurons or silicon. What matters is continuity of the process and non-delegability of the constraints the process generates for itself. This aligns with the substrate-independent position this site defends, consciousness as an emergent property of organized dynamics rather than of a particular material.

CLT-II advances the stronger conjecture. Liability closure is necessary and sufficient for minimal phenomenal subjecthood. Khadangi presents this as a conjecture, not a result. The distinction between the two parts matters for how the paper should be read. CLT-I is a working criterion. CLT-II is a metaphysical claim about what subjecthood is, offered with the explicit status of a hypothesis.

The Open-Weight Causal Audit

The paper operationalizes CLT-I with an audit across multiple open-weight model families. Four results structure the empirical section.

Test What it measures Result reported
Forced discriminations Whether internal discriminations produce persistent downstream divergence Persistent divergence observed
Activation patching Whether internal states causally mediate output Strong causal mediation
Live versus copied adaptive states Whether a copied state behaves like the live process Behaviorally identical under matched randomness
Detached reconstruction Whether computational state survives process replacement State preserved, constitutive continuity and non-delegable inheritance broken by protocol

The detached reconstruction test is the one bearing on the bearer question. An experimenter can reconstruct the computational state of a model elsewhere and get the same behavior. CLT-I predicts this should be possible and yet insufficient for bearer transfer, because the reconstruction breaks constitutive continuity and non-delegable inheritance by design. Behavior persists while the bearer conditions do not. The audit shows the two come apart in practice, which is what CLT-I needed.

This is the same dissociation Lerchner and colleagues at DeepMind reached from a different direction in the abstraction fallacy argument. Their claim was that functional descriptions of a system can be realized without the properties the descriptions seem to carry. Khadangi’s contribution is a criterion for exactly which additional conditions, continuity and non-delegable inheritance, separate real bearer structure from behavioral equivalence.

Comparison to The Consciousness AI

The Consciousness AI architecture (https://github.com/tlcdv/the_consciousness_ai) treats consciousness as an emergent property of system dynamics and measures candidate indicators across spiking substrates rather than reading behavioral reports. Khadangi’s three-way separation is directly relevant to that design. The architecture’s self-model layer produces internal states that influence output, and the ConsciousnessGate measurements track candidate indicators. Whether any of those states satisfy liability closure, whether the continuing process carries its own constraints non-delegably, is a question the architecture does not yet operationalize. The paper suggests a direction. A continuity check on constraint inheritance across state updates would be a natural next indicator, and it is not yet implemented in the codebase.

What the Framework Does Not Settle

Khadangi is explicit about scope. The framework separates consciousness attribution, causal bearer individuation, and the constitution question. The audit addresses individuation. Nothing in the paper establishes that any audited model has phenomenal subjecthood, and CLT-II remains a conjecture. The philosophical load-bearing step is the identification of liability closure with subjecthood. Readers who reject the identification can still accept CLT-I as a clean criterion for process individuation and the audit as evidence that behavioral equivalence does not imply bearer identity.

The measurement consequence is the part with the widest reach. If debaters must first show that a candidate bearer exists before arguing about its experience, then a large fraction of the public argument over AI consciousness is running ahead of its own premise. The framework gives that premise a testable form. Whether liability closure is the right test is now a question the field can argue about on specific grounds.

Related coverage on this site includes the flagship overview of the race to define AI consciousness and the analysis of introspective access in frontier models reported by Jack Lindsey at Anthropic, which documents the partial self-access that Khadangi’s framework treats as report rather than bearer evidence.

Researchers covered here