Fork the consciousness, or download the project and create your own.

Behavioral Lift Shows Reasoning Training Amplifies the Wrong Behaviors

Does reasoning-oriented training make models better reasoners, or does it just make their reasoning traces look better. Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, and Graham Neubig answer with a measurement framework in “Amplified Does Not Mean Predictive. Reasoning Behaviors in Thinking Models” (arXiv:2608.13760, posted 13 August 2026, published at COLM 2026). The paper introduces Behavioral Lift, a metric for how much correctness changes when a reasoning behavior is present in a model’s trace versus absent. Across 15 models, 6 benchmarks, and 15,282 annotated traces, the authors find a systematic mismatch they call the Amplification-Lift Gap. Training amplifies the behaviors that look deliberative. The behaviors that actually predict correct answers get amplified far less, and one of them is self-awareness.

The distinction the metric encodes is between amplification and lift. Amplification is how much more often a behavior appears in a thinking model’s traces. Lift is how much that behavior’s presence changes the probability that the answer is correct. Training changes the first quantity directly. The paper asks whether the second one moves with it.

The Amplification-Lift Gap

The study annotates 15,282 traces from 15 models across 6 benchmarks, spanning text-only and vision-language reasoning, with a taxonomy whose core behaviors are defined for both LLM and VLM traces. The pattern that emerges is consistent across modalities.

Behavior Amplified by training Associated with correctness
Self-correction Strongly Moderate
Hypothesis testing Strongly Moderate
Uncertainty acknowledgment Strongly, 3 to 7 times Weak or negative
Confidence calibration Barely Strongest positive signal in both modalities
Knowledge alignment Weakly Among the highest-lift behaviors
Self-awareness Barely Among the highest-lift behaviors

Two rows carry the paper’s argument. Uncertainty acknowledgment is the behavior training amplifies most aggressively, three to seven times, yet its association with correctness is weak or negative. A model that hedges more after reasoning training is not a model that answers more correctly. And confidence calibration, among the strongest positive signals of correctness in both text and vision-language traces, is barely amplified at all.

The hedging result is the sharpest case, and it deserves a careful reading. A trace that names its own uncertainty reads as careful reasoning, to a human evaluator and plausibly to a preference-trained model alike. The paper does not identify the training mechanism behind the amplification, and its authors are explicit that reasoning-oriented training was studied as a class rather than decomposed into specific objectives. What the numbers establish is the mismatch itself. The behaviors most rewarded with amplitude are not the behaviors most rewarded with correct answers, and the two currencies come apart by a factor of several in opposite directions. The authors’ proposed remedy, process-level objectives that reward calibrated and grounded reasoning rather than surface form, follows directly. So does the harder question of why surface form was being rewarded in the first place.

The conclusion the authors draw is that reasoning-oriented training does not preferentially amplify the highest-lift behaviors. The traces look more deliberative. The deliberation is not concentrated where it pays.

Self-awareness as a measured behavior

The result that matters most for consciousness research is the treatment of self-awareness as a measurable, correctness-relevant behavior rather than a self-report. In this taxonomy, self-awareness is a trace-level pattern, something a model does in the course of reasoning, scored by annotators and validated against outcomes. On that operationalization, self-awareness is among the highest-lift behaviors in the study. Models whose traces show it are more likely to be right.

That operationalization is the same move the site’s indicator discipline calls for, separating behavioral function from experiential claim. Keith Frankish’s illusionist position on LLM first-person reports argues that reports of inner states are outputs like any other. Behavioral Lift goes further in a useful direction. It scores the behavior without asking for a report at all, and it prices the behavior in the only currency the benchmark can measure, correctness. Whether trace-level self-awareness has any connection to experience is untouched by the paper. What the paper establishes is that the functional version of the property is real, measurable, and valuable to the system.

The result also sharpens a question the jacobian lens work on Claude’s internal representations raised from the other direction. That research found privileged internal representations bearing functional hallmarks of conscious access. Behavioral Lift finds behavioral markers of self-awareness that training leaves mostly unamplified. Both lines locate something self-directed in current systems that standard training does not select for, one inside the weights and one in the traces. Whether the two markers track the same underlying capacity is an open question neither paper answers.

Comparison to The Consciousness AI

The Consciousness AI project’s documented architecture computes a global workspace signal from causal gate states and shapes behavior through an affective core, as described in the project architecture overview. The project documentation does not describe a trace-level behavioral audit of its agents, and no Behavioral Lift analysis of the project’s own runs exists. The transferable idea is procedural. The lift metric is model-agnostic, and the project’s agents could be audited the same way, annotating traces for calibration and self-awareness behaviors and testing which of them predict task success. That audit is this analysis, not the paper’s, and it remains untested.

The state of the field review records that no current AI system meets the behavioral indicators of consciousness. The present paper adds a subtler point to that record. The behaviors nearest to self-directedness that current models do show are exactly the ones the training pipeline does not select for, which means the gap between what models can do and what they are trained to display is measurable, and currently favors the display.

Limits

Behavioral Lift is a correlational measure over annotated traces, so lift values describe associations within this benchmark suite, and the annotation taxonomy is one design among several. Amplification was measured against reasoning-oriented training in general rather than against specific training objectives, so the paper identifies the mismatch without isolating its mechanism. The authors’ proposed remedy, process-level objectives that reward calibrated and grounded reasoning rather than surface form, is a research direction, not a demonstrated fix. What is established is that the mismatch exists at scale, that it survives across modalities, and that the behaviors worth amplifying can be named and priced.