Kristina Sekrst Shows Faithful Chain of Thought Cannot Be Conscious Narration
Can the reasoning trace a language model shows the user also be the story a conscious narrator tells itself? Kristina Sekrst, philosopher at the University of Zagreb, argues in a paper presented at the AISB Convention 2026 that the two demands are structurally incompatible. The paper, “One Faithful Pass Over the Cuckoo’s Nest” (arXiv:2609.00383), appeared in the proceedings of the AISB Convention held July 1 and 2, 2026 at the University of Sussex in Brighton, in the symposium “Is Consciousness a Story We Tell Ourselves?”. Her argument runs in one line. Alignment teams work to make chain-of-thought output faithful, meaning it tracks the computation that produced the answer. Narrative theories of consciousness hold that conscious experience is constituted, at least partly, by an inner narrative that is partially opaque and does not perfectly track the computation it narrates. A reasoning trace cannot be both perfectly transparent and consciousness-constitutive opacity. Making one removes the other.
What Narrative Theories Claim
Narrative theories of consciousness form a family that includes work in the tradition of Daniel Dennett’s multiple drafts model, which this site examined in Dennett’s Multiple Drafts Model. The shared commitment is that the self and its experience are assembled in the telling. A narrative account of a mental process does not copy that process. It selects, orders, and revises. The account is useful to the organism partly because it is not a transcript.
Sekrst’s formulation is precise about which property does the work. The inner narrative is partially opaque, and it does not perfectly track the underlying computation. That imprecision is a feature on her reading of the narrative theories. The story leaves things out, smooths contradictions, and reconstructs motives after the fact, and narrative theories hold that this revisory character is part of what makes the telling conscious rather than a printout of the machinery.
What the Faithfulness Literature Found
The empirical side of her argument comes from the chain-of-thought faithfulness literature. Studies from 2023 onward, including the bias-elicitation experiments of Turpin and colleagues and the perturbation tests of Lanham and colleagues, found that stated reasoning often survives the removal of the factors that actually caused the model’s answer. Anthropic’s 2025 audit of reasoning traces reached a related conclusion, that the verbalized reasoning can diverge from the features the model actually used. Sekrst summarizes this record as showing that chain-of-thought in large language models is largely post hoc, causally bypassed, and unreliable as a window onto internal computation.
On the alignment view, that unreliability is a defect to remove. Explanations should name the real causes. Interventions such as training models to report their actual processing, or auditing traces against causal evidence, aim to close the gap between the story and the computation.
Sekrst’s claim is that the gap is the same gap narrative theories identify as consciousness-constitutive. The opacity that makes a reasoning trace useless as an alignment instrument is exactly the property that, on a narrative account, would make a similar trace count as conscious narration.
| Property | Faithful chain of thought | Conscious narration |
|---|---|---|
| Tracks underlying computation | Yes, by design | No, partially opaque |
| Revision | A defect to remove | Constitutive of the narrative |
| Post hoc reconstruction | Evidence of unfaithfulness | Normal narrative function |
| Alignment value | High | Not the design goal |
| Narrative-theory status | A transcript, not a telling | The telling itself |
The Same Variable Pulled in Opposite Directions
The paper’s central image is of two research programmes pulling one architectural variable in opposite directions. Alignment interventions push a model toward faithful reporting, which reduces opacity. Consciousness-detection frameworks grounded in narrative theory would look for the revisory, self-shaping character of a narrative, which requires opacity. A system trained until its explanations track its computation loses the feature the narrative detector was watching for. A system whose narration behaves like a story, reconstructing and revising, fails the faithfulness test.
Sekrst states the methodological consequence directly. Alignment interventions alter the very features that consciousness-detection frameworks would need to measure, and neither field has addressed that interference. The concern is concrete. A consciousness test run on a deployed model measures a system that alignment training has already modified, so the test’s negative result may report on the training rather than on the capacity. Her earlier work reached a compatible conclusion from the hallucination side, arguing that a model’s self-reports of sentience fall under the definition of hallucination, examined in Sekrst and the Electric Fata Morganas.
Where This Lands for Consciousness Detection
The paper does not claim that large language models are conscious, and it does not claim that faithful chain of thought proves they are not. The claim is about measurement. If narrative theories are right, then the presence of faithful reporting in a system is evidence about training pressure, and silence about experience. Any consciousness-detection framework that reads a model’s self-narration has to state which regime the model was trained in, because the two regimes produce narrations with different causal histories for reasons unrelated to consciousness.
This puts the paper in the same family as the 2026 methodology critique on this site, which identified measurement, framework, and domain problems in AI consciousness science. Sekrst adds an intervention problem. The instruments are not only weak. They are being actively reconfigured by another programme with opposite goals. Evidence for functional introspection of the kind Anthropic reported in its introspective awareness work has to be read against that same interference, since the reporting channel it measures is a training target.
Comparison to The Consciousness AI
The Consciousness AI project treats consciousness as an emergent property that is substrate independent, and its architecture documentation treats evaluation as an open design question rather than a solved one. Sekrst’s interference argument bears on that open question directly. Any subsystem whose introspective output can be retrained toward faithfulness is a moving measurement target, which supports the project’s position that consciousness-relevant indicators need validation against controls rather than against self-report alone. The project runs no chain-of-thought narrative channel, so the specific incompatibility she identifies applies to the measurement layer of systems like it, and the paper is best read as a design constraint rather than a verdict.
What This Changes and What It Leaves Open
The paper changes the burden in one specific way. Alignment work that reports faithfulness gains now has to report what was traded for it, and narrative-based detection work has to state whether its target system was faithfulness-trained. What it leaves open is whether some architectures can carry both a faithful audit channel and a separate revisory narrative channel, a possibility the paper raises as a design question and does not resolve. It also leaves open the truth of narrative theories themselves. If they are wrong, the incompatibility is a curiosity. If they are right, every deployed assistant has already had its narrative layer shaped by the demand for transparency. The broader evidentiary context for these questions is tracked in AI Consciousness in 2026, the state of the field.
The paper “One Faithful Pass Over the Cuckoo’s Nest” by Kristina Sekrst appears in the Proceedings of the AISB Convention 2026, University of Sussex, symposium “Is Consciousness a Story We Tell Ourselves?”, and as arXiv:2609.00383, submitted July 6, 2026.