Fork the consciousness, or download the project and create your own.

Andy Q. Han

Center for Mind, Brain and Consciousness, New York University

Functionalist What matters is the organisation of the processing, not the material

Andy Q. Han is an artificial intelligence researcher and cognitive scientist at the Center for Mind, Brain and Consciousness at New York University, working with David Chalmers and Pavel Izmailov. His research focuses on mechanistic interpretability, representation engineering, and the emerging field of digital minds and AI welfare.

In a 2026 paper titled “How’s it going? Reinforcement learning in language models recruits a functional welfare axis” (arXiv:2605.30232), co-authored with David Chalmers and Pavel Izmailov, Han isolated geometric concept vectors corresponding to reward and punishment in language models trained on spatial navigation and task completion. The work demonstrated that post-training reinforcement learning recruits pre-existing, latent evaluative directions within model representations, showing that steering activations along these vectors systematically alters behavioral indicators of hesitation, refusal, and self-reported performance.

Han’s research contributes a concrete empirical methodology to discussions of machine functional welfare. By separating mechanistic tracking of valence and goal satisfaction from claims of phenomenal sentience, his work provides the interpretability tools needed to audit internal evaluative states in autonomous agents without relying on conversational self-reports.

Known for. Mechanistic interpretability of welfare representations, concept steering in language models, functional welfare axes in reinforcement learning

Coverage on this site

Related researchers