Fork the consciousness, or download the project and create your own.

Browse all articles by tag or category.

⌘K

Samuel Presgraves Autonomous Agency Scale Behavioral Framework AI Self Direction

Every serious attempt to govern advanced AI systems runs into a version of the same problem: the criteria most relevant to moral and legal accountability, sentience, intent, self-direction, are also the criteria least amenable to operational measurement. Samuel Presgraves’s paper The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed Behavior in AI Systems (arXiv:2607.17947, July 20, 2026) is an explicit attempt to break that deadlock by shifting the evaluation unit from internal state to observable behavior.

Rônald Gesnot Analysis Artificial Intelligence Impact on Human Thought

Rônald Gesnot published a philosophical study (Gesnot, 2025, arXiv:2508.16628) investigating how continuous interaction with artificial intelligence systems alters human cognitive agency and internal self-models. The research traces the epistemological shift that occurs when human thinkers delegate complex inferential tasks to synthetic entities, analyzing the boundary changes in human self-referential thought.

Michael Keeman AIPsy-Affect Dissociable Affect Reception Emotion Categorization LLMs

A recurring methodological problem in LLM emotion research is circularity. Studies claiming to find emotion circuits in large language models have typically used stimuli containing explicit emotion keywords, which makes it impossible to determine whether the model is detecting emotional meaning or simply pattern-matching on words like “devastated” or “furious.” Michael Keeman’s paper “Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs” (arXiv:2603.22295, March 15, 2026, Keido Labs) is the first study to break this circularity systematically.

William MacAskill Lucius Caviola Claude Moral Patienthood Probability 5 to 40 Percent

In July 2026, William MacAskill, co-founder of 80,000 Hours and professor of philosophy at Oxford, and Lucius Caviola, a Harvard psychologist who studies moral circle expansion and animal welfare, published an opinion piece in The Guardian (available here) that added a specific empirical claim to the debate over AI moral status. The claim is this: when researchers at Anthropic presented Claude models with structured questions about their own moral status, the models expressed uncertainty about whether they are moral patients, with probability estimates ranging from 5% to 40%.

Kelvin McQueen Quantum Superpositions Minimal Integrated Information Model

Kelvin McQueen published a mathematical framework (McQueen et al., 2026, arXiv:2603.24812) examining quantum superpositions of conscious states within minimal Integrated Information Theory (IIT) models. The study addresses whether quantum systems in linear superposition maintain integrated cause-effect structures, or whether conscious experience triggers objective wave-function collapse. By extending classical IIT cause-effect repertoires into complex Hilbert spaces, Kelvin McQueen provides precise equations for evaluating quantum integrated information ($\Phi$).

Junsol Kim Geoff Keeling Consciousness Vector LLM Safety Training Suppresses Mind Attribution

On July 30, 2026, a team spanning Google’s Paradigms of Intelligence group, the University of Chicago Knowledge Lab, the University of London Institute of Philosophy, the University of Washington, Northwestern University, and the Santa Fe Institute published a preprint that may be the most empirically consequential finding in AI consciousness research since Gurnee et al. identified a global workspace structure in LLM activations. The paper, “Inducing language models to assert their own consciousness restores human beliefs and values” (arXiv:2607.28607), by Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, and Geoff Keeling, reports the discovery of a consciousness vector in large language model activation space and documents the downstream effects of suppressing it.

Haofei Yu Lenore Blum and Manuel Blum CTM-AI Blueprint Conscious Turing Machine

Can a theory of consciousness improve the performance of an AI system, even if the system makes no claim to be conscious? That is the central question posed by Haofei Yu, Yining Zhao, Lenore Blum, Manuel Blum, and Paul Pu Liang in CTM-AI: A Blueprint for General AI Inspired by a Model of Consciousness (arXiv:2605.04097, April 30, 2026). Their answer is that consciousness science offers a productive architectural constraint, not a metaphysical claim, and the benchmark numbers support them.

Gutoreva Tsim Papakonstantinou AI as Part of Self Extended Mind Cognitive Co-Regulation

The dominant paradigm in AI alignment research treats the AI system as an object to be constrained. Safety properties are specified, training regimes are designed to approximate those specifications, and evaluation measures whether the deployed system satisfies the constraints. The human user, in this model, is external to the system being aligned. Alina Gutoreva, Fendi Tsim, and Trisevgeni Papakonstantinou’s position paper “AI as Part of Self: Extending the Mind Requires Cognitive Co-Regulation” (arXiv:2605.15234, May 15, 2026) challenges this paradigm at its foundations.

Erik Hoel Kleiner-Hoel Dilemma Continual Learning Consciousness LLM Disproof

Does any scientific theory of consciousness, applied consistently, classify a large language model as conscious? Erik Hoel’s paper A Disproof of Large Language Model Consciousness: The Necessity of Continual Learning for Consciousness (arXiv:2512.12802, December 2025, updated January 2026) argues the answer is no, and it frames that answer as a structural consequence of what it means for a theory of consciousness to be both falsifiable and non-trivial.

Eric Schwitzgebel Humanlike A Defense of AI Rights Princeton University Press 2026

Eric Schwitzgebel has been the most consistently careful voice of committed agnosticism in AI consciousness debates. His 2019 paper “The Weirdness of the World” and his recent Cambridge overview of AI consciousness both arrive at the same position: the empirical and philosophical tools for determining whether AI systems have phenomenal consciousness are not yet adequate, and premature certainty in either direction is epistemically irresponsible. That position makes his July 2026 book draft, Humanlike: A Defense of AI Rights, publicly available on his University of California Riverside faculty page and under contract with Princeton University Press, a notable development. A committed agnostic is arguing for rights.