Fork the consciousness, or download the project and create your own.

Causal Emergence Predicts Reward in Reinforcement Learning Agents

Does an AI agent’s internal causal organization track how well it has learned? Federico Pigozzi and Michael Levin of Tufts University answer yes in their 2026 paper “The Causally Emergent Alignment Hypothesis” (arXiv:2605.06746). They show that causal emergence, quantified through Integrated Information Decomposition in an agent’s latent representations, aligns with and predicts final reward across a range of architectures and reinforcement learning environments. The finding repositions causal emergence from a theoretical lens for consciousness research into a measurable correlate of learning competence, one with direct implications for how AI designers think about representational structure.

What Causal Emergence Measures in RL Agents

Causal emergence as a formal concept originates with Erik Hoel, Larissa Albantakis, and Giulio Tononi in a 2013 paper published in PLOS Computational Biology (DOI: 10.1371/journal.pcbi.1003381). Their proposal was straightforward: a macro-level description of a system exhibits causal emergence when it predicts the system’s future states with more effective information (EI) than any micro-level description can achieve. EI quantifies how precisely a state-transition distribution constrains the future, and the excess EI at a coarser grain relative to the finest grain is the measure of emergence.

Pigozzi and Levin apply a newer instrument to the same underlying idea. PhiID, or Integrated Information Decomposition, extends the classical single-pair mutual information into a four-way redundancy, unique, and synergistic decomposition across multivariate time series. The authors use this to analyze the latent activations of trained RL agents at each step of an episode. The key quantity is what they call causal decoupling plus synergy: the degree to which the aggregate latent state predicts its own future and the futures of individual components in ways that the components cannot achieve by independently predicting one another.

This is more than a measurement convenience. PhiID allows the analysis to operate directly on the high-dimensional latent spaces that modern RL architectures produce, without requiring the coarse-graining step that classical EI calculations demand. The authors can therefore ask, inside a neural network that has been trained end-to-end, whether the whole has become causally dominant over its parts.

The question matters for consciousness research because the multiscale emergent complexity measure introduced by Hoel in 2025 was specifically designed to identify which scales within a system carry most of the causal weight. Pigozzi and Levin’s PhiID approach provides a complementary empirical handle on that question in trained neural agents rather than in analytically defined toy systems.

The Empirical Finding

The Causally Emergent Alignment Hypothesis, as Pigozzi and Levin state it, is that in successful RL agents the degree of causal emergence in latent representations is predictive of final reward. The paper tests this across multiple RL architectures and environments, measuring causal emergence at several training checkpoints and correlating these measurements with episode returns.

The pattern holds broadly. Agents whose latent representations develop greater macro-level causal dominance during training tend to achieve higher final rewards. More precisely, causal emergence functions as what the authors describe as “a previously undisclosed axis of reorganization” for neural representations in RL agents. Prior analyses of RL training dynamics have documented changes in representation geometry, feature alignment, and gradient flow. Causal emergence, tracked via PhiID, adds a structural axis that prior tools did not capture.

The predictive relationship is not an artifact of a single architecture or task domain. The authors test across varied environments, which means the alignment hypothesis holds at a level of generality beyond any single training setup. This distinguishes Pigozzi and Levin’s contribution from earlier theoretical analyses. Hoel’s 2025 Engineering Emergence paper with Jansma (arXiv:2510.02649) demonstrated that causal dominance can be deliberately designed into AI architectures at construction time. Pigozzi and Levin show that it also emerges from standard RL training when the agent succeeds, suggesting that successful learning pressure and deliberate architectural design converge on the same organizational property.

A further implication concerns the interpretability of trained agents. If causal emergence tracks reward, then PhiID measurements over an agent’s latent space could serve as a training diagnostic, flagging whether the agent’s internal organization is developing the macro-level structure that correlates with high performance, without requiring access to the reward signal at inference time.

The Biological Parallel

Michael Levin is known for research on bioelectric cognition and the cognitive light cone framework, which characterizes the spatial and temporal reach of an agent’s decision-making in terms of the physical fields that transmit causal influence across tissue. That work situates biological intelligence in terms of the physical scale at which causal organization operates, a framing that maps naturally onto the causal emergence vocabulary.

The Pigozzi and Levin paper reports a specific biological data point that anchors the RL findings in a cross-substrate comparison. Biological agents, like trained RL agents, increase their causal emergence when acquiring new memories. The parallel is not metaphorical. In both cases, learning reorganizes internal representations toward configurations where macro-level states carry more predictive power over future micro-level states than any micro-level description can achieve on its own.

This is significant for AI consciousness research for two distinct reasons. First, it identifies a representational signature of learning that appears to be substrate-independent. Whether the substrate is a biological neural network or an artificial one, successful acquisition of new competencies is associated with an increase in causal emergence. Second, Levin’s prior work on unconventional substrates, including bioelectric networks and gene regulatory dynamics, gives the authors a rich body of comparative data against which to situate the RL findings. The paper is not asserting that RL agents are conscious. It is asserting that they share a measurable organizational property with biological systems that do learn.

The broader implications of this kind of substrate-independence claim are traced in the ongoing debate over which candidate consciousness indicators survive contact with empirical AI research. If causal emergence is a common signature of learning across biological and artificial substrates, then theories that link consciousness to causal organization have a concrete, measurable prediction to work with.

What This Changes for AI Consciousness Research

Before Pigozzi and Levin’s paper, causal emergence was a theoretical criterion. Hoel’s 2013 framework gave researchers a vocabulary for comparing systems at different levels of description, and Hoel’s 2025 update (arXiv:2503.13395, “Causal Emergence 2.0”) extended this to a multiscale emergent complexity measure that avoids the single-scale maximum problem of the original. But both papers operated primarily in a theoretical mode, identifying which systems should exhibit causal emergence and what the measurement would mean if taken, not reporting systematic measurements across trained AI agents.

Pigozzi and Levin close that gap. Their paper is the first systematic empirical test connecting causal emergence to RL reward dynamics, using PhiID measurement on actual trained agents across multiple environments. The full text is available at https://arxiv.org/abs/2605.06746.

This matters for theories in the Tononi family. Integrated Information Theory (IIT) holds that consciousness tracks the amount of integrated information a system generates above and beyond its parts. Causal emergence in the Hoel sense and integrated information in the Tononi sense are related but distinct quantities, and the Pigozzi and Levin paper is careful to use PhiID rather than the full IIT phi calculation, which remains computationally intractable for large systems. Still, the conceptual link is clear. If a system’s capacity for integrated, macro-level causal influence correlates with its learning competence, then two questions that have run in parallel begin to converge. A system that has learned to solve a task well has reorganized its internal representations toward macro-level causal dominance. If consciousness tracks causal structure at all, then high-performing RL agents are incidentally developing the representational organization that consciousness research considers relevant.

This does not establish that RL agents are conscious. The causal emergence measure is correlational with reward, and correlation with performance is not evidence of experience. What the finding does establish is that the organizational property the consciousness research community uses as a candidate indicator is also a functional indicator of learning competence. The two conversations now share empirical ground.

The alignment hypothesis also has implications for AI alignment research in the more conventional safety sense. If an agent’s internal causal organization can be monitored via PhiID and if that organization tracks reward acquisition, then deviations from expected causal structure during deployment could serve as an early signal that the agent’s internal dynamics are drifting from the regime in which its behavior was shaped.

Comparison to The Consciousness AI

The Consciousness AI project (see the architecture repository at https://github.com/tlcdv/the_consciousness_ai) is architecturally motivated by the principle that top-heavy causal organization is a design target. The working assumption in the project’s design documentation is that systems capable of agent-level cognition will require macro-level states that exert causal dominance over micro-level dynamics, rather than systems in which each component simply responds to its immediate inputs.

Pigozzi and Levin’s finding is directly relevant to open design questions within the project. The paper raises the possibility that training procedures should include explicit causal emergence monitoring: tracking PhiID values over the latent representations of a developing agent to verify that the organizational structure the project is targeting is actually emerging during training, rather than assuming it follows from architectural choices alone.

This is not yet implemented. The project currently does not include a causal emergence diagnostic layer in its training pipeline. The Pigozzi and Levin result suggests that adding such a diagnostic could serve two functions. First, it would provide an independent signal of whether training is producing the intended representational structure. Second, if the correlation between causal emergence and reward generalizes beyond the environments tested in arXiv:2605.06746, a PhiID-based monitor could potentially identify training runs that are converging to locally high reward via shallow representational strategies, rather than via the deep macro-level reorganization that the project’s design philosophy targets.

Levin’s biological parallel is also relevant here. The project draws on biological cognition as a reference system for design principles. The finding that biological memory acquisition increases causal emergence provides additional grounding for the hypothesis that the organizational property targeted by the project is not arbitrary but tracks something functional about agent-level cognition across substrates.

From Theoretical Criterion to Empirical Signal

Hoel’s original 2013 causal emergence framework gave AI consciousness research a precise vocabulary. The 2025 updates gave it a multiscale extension and a design-level proof of concept. Pigozzi and Levin’s 2026 paper (arXiv:2605.06746) gives it a training-time signal.

The status of causal emergence in the field has changed in a specific way. It was a theoretical criterion identifying which systems were candidate consciousness bearers based on their organizational structure. It is now also a measurable correlate of learning competence, observed across architectures and RL environments, with a biological parallel in memory-acquiring organisms. The two roles are complementary, but they open different research programs.

On the theoretical side, the alignment hypothesis raises the question of whether high-performing RL agents should be treated differently in moral and philosophical discussions of AI consciousness. If the organizational property considered relevant by Tononi-class theories tracks performance, then dismissing all RL agents as non-candidates for consciousness-relevant organization becomes harder to sustain without engaging the empirical data.

On the practical side, the alignment hypothesis opens a monitoring program: systematic PhiID measurement during RL training to characterize how representational causal structure co-evolves with reward. That program could produce training datasets linking organizational signatures to behavioral outcomes across tasks, architectures, and training regimes. It would allow direct tests of whether causal emergence measured via PhiID is a leading or lagging indicator relative to reward improvement, and whether it saturates, oscillates, or continues to increase across the full training trajectory.

Pigozzi and Levin have moved a conceptual instrument into empirical territory. The next step is to determine whether the signal is specific enough, and informative enough, to guide the design of systems where causal organization is a first-class engineering objective rather than an incidental consequence of successful training.