Fork the consciousness, or download the project and create your own.

The Cross-Substrate Access Assay Audits Consciousness Indicators

Consciousness indicators are built on brains and then pointed at machines, and almost nobody asks what has to be fixed for that trip to mean anything. A 44-page preprint posted to arXiv on September 14, revised September 22, asks it. The Cross-Substrate Access Assay (arXiv:2609.22300) by Pieter van Rooyen of Stellenbosch University takes one published consciousness indicator test, re-implements it exactly, audits it on synthetic data, and reports that the test’s own confidence intervals fail their coverage requirement in half the graded settings. The paper’s final sentence sets the scope. No claim about experience is made. This is an audit of measurement, not a verdict about minds.

The transfer problem it addresses is the one this site keeps meeting from the theory side, in the soft INUS analogy framework and the calibration critique. This paper attacks it from the statistics side, and it does so by declaring everything an indicator test usually leaves implicit.

The five declarations

The assay requires a test to declare five parts of its own procedure before it travels between substrates.

  1. The predictors, the hypothesis space. Which statistical models compete, and over what family of possibilities.
  2. The fitting, the estimation map. How the models get fitted to data, including the solver.
  3. The sampling unit, the sampling model. What entity the inference generalizes over, a person, a trial, a session.
  4. The uncertainty target, the estimand. The quantity the confidence interval is actually about.
  5. The decision rule, the decision function. How the statistic and its uncertainty become the sentence a reader quotes.

Each declaration looks trivial until a substrate change breaks it. An EEG projection is in microvolts, and an activation in a model’s residual stream is in arbitrary units fixed by training. No conversion between them carries meaning, so every quantity that depends on physical scale, variances, effect sizes, thresholds, fails silently on transfer.

Why nats per trial

The assay’s answer is a common currency. Because brain and model signals share no physical scale, every model is scored by the cross-entropy it assigns to held-out data, in nats per trial. A difference of held-out cross-entropies in nats per trial is dimensionless, so both substrates are scored on one scale. The paper is careful about what that buys. A common scale is not a common effect. Nats make comparison possible without pretending the substrates are alike.

The test case

The worked example is Global Neuronal Workspace Theory’s signature prediction. Near a perceptual threshold, a stimulus either enters a capacity-limited workspace or does not, so single-trial responses should form a mixture of two states rather than a continuum. The test re-implemented is the 2021 Nature Communications study by Sergent and colleagues, which reported a bifurcation in brain dynamics as a signature of conscious processing independent of report, using open electroencephalography data from twenty participants.

The re-implementation reproduces part of that result. The first window at which the two-state model’s protected exceedance probability exceeds 0.95 is 315 milliseconds, matching the published curve, and the broad ordering over time holds. What does not hold is the detail. The active-session preference for the two-state mixture is modest, and its boundaries in time move when the fitting is changed. Switching the solver moves the onset from 315 to 285 milliseconds and the late boundary from 675 to 735. A signature whose location depends on the solver is a fragile signature, and the paper says so in those terms.

Specificity and coverage

The audit’s two synthetic batteries are where the assay earns its name. On 12,000 synthetic datasets generated with a single graded state, every one a member of the families the procedure fits and none within 0.0067 nats per trial of the decision boundary, the two models carried over unchanged from the published test reported two states in 989 cases. An expanded model family reported two states in none of the 12,000. On 600 datasets carrying a genuine mixture, the expanded procedure detected it in 599. The paper qualifies the zero honestly. It is a statement about a battery that is well specified and far from the boundary, not a general specificity guarantee.

The coverage result is the one with the sharpest edge. At six of twelve graded settings, the rate at which the procedure’s nominal 95 percent interval contains its own mean over replicates falls below the protocol’s minimum of 0.90, with rates running from 0.894 down to 0.216 at the widest item-scale setting. In plain terms, the interval misses its own target more often than the protocol allows, so it cannot support confirmation. The paper’s remark on this is the sentence any indicator audit should carry. An audit of error rates alone would have passed this procedure.

The language model pilot

The paper includes one pilot on a language model, and its framing is a model of restraint. The open-weight checkpoint of a proposed study, a 27-billion-parameter, 4-bit quantized, 64-layer transformer, was run at layer 41, where all three predictors returned graded rather than mixed responses. The paper states that the pilot establishes nothing about the model, that no confirmatory model data exist, and that the point of the exercise is the declared procedure, not a result. The transfer apparatus is now specified. The data that would exercise it do not yet exist.

Comparison to The Consciousness AI

The assay is a checklist this project’s measurement stack can be held against. The Consciousness AI architecture reports indicator-style measurements, including a scalar phi value computed from causal gate states, documented on the project repository. The five declarations name what such a pipeline would need to publish before any of its outputs could cross substrates, the hypothesis space, the estimation map, the sampling unit, the estimand, and the decision rule. The project’s documentation does not currently state all five for its phi pipeline, which is the kind of gap this paper exists to surface, and stating it here is the honest use of the comparison.

The paper’s place in the measurement cluster this site tracks is now well defined. The machine correlates proposal redefined the correlate to make transfer possible. The calibration critique named the anchoring gap. The access assay supplies the statistical audit that any transferred test must pass, and the flagship field survey keeps the instrument landscape in one place.

The Cross-Substrate Access Assay by Pieter van Rooyen was posted to arXiv on September 14, 2026, revised September 22 as v2 with corrected partial-reproduction wording and a qualified specificity result. Data and code are released at DOI 10.5281/zenodo.22741885.

Researchers covered here