Fork the consciousness, or download the project and create your own.

Larissa Albantakis and Toy Models for Testing Theories of Consciousness

Testing a theory of consciousness against a system as complex as a human brain or a large language model means testing it against something whose internal workings are only partially known, which makes it hard to say with confidence which features of the system the theory’s prediction actually depends on. Larissa Albantakis’s answer, developed in “On the utility of toy models for theories of consciousness” (arXiv:2508.00190, forthcoming as a chapter in Springer’s The Scientific Study of Consciousness), is to deliberately go small instead, using systems simple enough to fully specify, where a theory’s prediction can be traced to an exact mechanism rather than left as an inference about an opaque system.

Approach What it controls for What it sacrifices
Testing on brains or large AI systems Ecological relevance, real cognitive complexity Full knowledge of the system’s actual causal structure
Testing on toy models Complete specification of every causal dependency Ecological relevance, since toy systems do not do anything cognitively interesting
Comparing IIT and GWT on the same toy model A shared, fully known ground truth both theories must predict against Nothing, if the model is chosen to be genuinely diagnostic rather than trivial

Why Complex Systems Make Poor Test Cases

Evaluating whether Integrated Information Theory or Global Workspace Theory correctly identifies which systems are conscious usually means applying both to something biologically or computationally complex, and checking whether their verdicts match intuitions about that system. Albantakis’s toy models argument is that this approach confounds two separate questions, whether a theory’s formal machinery gives a determinate answer at all, and whether that answer is correct, because in a sufficiently complex system it is often unclear which part of the theory’s calculation is doing the work.

A toy model, a small system of logic gates or simple nodes with fully specified connections and update rules, removes that ambiguity. Every causal dependency in the system is known by construction, so when IIT computes a phi value or Global Workspace Theory identifies a broadcasting bottleneck, the result can be traced exactly to the specific structural feature responsible, rather than left as an opaque output of a black-box calculation. This lets researchers isolate exactly which architectural feature IIT or GWT is actually responding to, information integration in one case, a particular routing bottleneck in the other, and check whether that feature is the one intuitively relevant to consciousness or an artifact of the theory’s formalism.

Where IIT and GWT Diverge on Identical Systems

Albantakis’s chapter builds toy models specifically designed to pull IIT and GWT apart, constructing systems where the two theories’ formal criteria are satisfied to different degrees, high integration with no broadcasting bottleneck, or a clear bottleneck with minimal integration, and asking what each theory is committed to saying about the result. The value of this exercise is that it exposes exactly where the theories make different predictions rather than leaving the disagreement at the level of verbal description, which is often where debates between IIT and GWT proponents stall. A toy system with high phi but no global broadcast mechanism forces IIT to attribute consciousness where GWT would not, and a toy system with a clean broadcast bottleneck but low integration does the reverse, giving researchers a minimal, fully transparent case study for each direction of disagreement.

The Cogitate consortium’s adversarial test of IIT against GWT in human subjects, covered in neither theory survived what the Cogitate consortium’s adversarial test found, found genuine empirical friction between the theories’ predictions using neuroimaging in biological brains. Albantakis’s toy model approach is the complementary, fully controlled counterpart to that result, isolating the same theoretical tension in systems small enough that there is no question of which feature is responsible for a given theory’s verdict.

Consequences for Evaluating Toy and Small AI Architectures

The direct relevance to artificial systems is that many proposed AI consciousness indicators are themselves tested first on small, simplified architectures before any claim is made about larger models, and Albantakis’s framework supplies the methodology for doing that testing rigorously rather than impressionistically. Her earlier work, co-authored as part of the Findlay et al. team distinguishing artificial intelligence from artificial consciousness, already established that behavioral competence and integrated causal structure can diverge sharply in engineered systems, covered in dissociating AI from artificial consciousness. The toy models chapter extends that lesson methodologically, showing how to build minimal systems that make the divergence between competing theories, not just between behavior and structure, fully explicit and reproducible.

Pedro Mediano’s integrated information decomposition work, covered in Pedro Mediano on integrated information decomposition and synergy, supplies one of the computational tools that would make applying this kind of controlled comparison to larger, partially opaque systems more tractable, decomposing an integration measure into interpretable components rather than a single aggregate number.

Comparison to The Consciousness AI

This project’s pre-registered comparison of AKOrN oscillatory binding against IIT phi, described on the architecture page, is closer in spirit to Albantakis’s toy model methodology than to testing on an opaque, fully trained system, since the architecture’s causal gate states are directly inspectable rather than inferred. The finding that oscillatory binding and integrated information were mechanistically opposed across four tested architectural variants, with no configuration reaching the pre-registered correlation threshold, is exactly the kind of clean, traceable result Albantakis’s framework is designed to produce, a divergence between two theoretically motivated measures that can be attributed to a specific structural feature of the tested systems rather than left as an unexplained discrepancy. The architecture itself is not a toy model in her narrow sense, since it retains considerably more moving parts than a minimal logic-gate system, but the project’s practice of running pre-registered, fully specified comparisons before drawing conclusions follows the same underlying discipline her chapter argues for.

What Toy Models Cannot Settle

Albantakis is explicit that toy models cannot establish which theory of consciousness is true, only which theory’s formal criteria are satisfied by which structural features, and how sharply those criteria diverge when tested against systems simple enough to leave nothing hidden. Whether phi or global broadcasting availability is the property that actually matters for experience is a question toy models cannot answer on their own, since a toy system’s simplicity is also why nobody would claim it is conscious regardless of what either theory computes. What her approach buys the field is a discipline for testing internal theoretical consistency and mutual compatibility before those same criteria get applied, with far less certainty about what is driving the result, to a brain or a large AI system. The wider landscape of theory comparisons this feeds into is tracked on the state of the field on AI consciousness.