Fork the consciousness, or download the project and create your own.

DOSE-I Records 1,129 Consciousness Transitions for Public Study

Consciousness indicators are easy to propose and hard to validate, because validated ground truth about conscious state is scarce. DOSE-I is a step against that scarcity. The dataset, “DOSE-I, A Multimodal Biosignal Dataset of Procedural Sedation for Endoscopy” (arXiv:2606.02570, posted June 30, 2026), by Jakob Garbe, Jan W. Kantelhardt, Katja Seeliger, and Thomas Schmid, publishes 78.5 hours of multimodal biosignal recordings from 171 sedation records. Within those recordings sit 1,129 annotated consciousness transitions, the moments a patient moves between conscious and unconscious states, and 7,328 sedation depth labels. The primary citation is the AIME 2026 conference, with preprocessing code released alongside the data on Zenodo.

What the Dataset Contains

Each record pairs biosignals with clinical annotations at two granularities. Sedation depth labels assign a level to each segment, 7,328 of them in total. Consciousness transitions mark the boundaries between levels, 1,129 of them. The signals are multimodal, which means several physiological channels are recorded in parallel over the same timeline, so models can learn what a transition looks like across channels rather than in one stream.

Item Count
Recording hours 78.5
Sedation records 171
Consciousness transitions 1,129
Sedation depth labels 7,328

Why Consciousness Transitions Are the Valuable Part

A consciousness indicator earns trust by passing a specific test. It should change when conscious state changes, and stay stable when conscious state does not. Testing that requires data where the state actually changed at known moments. Healthy waking subjects offer almost none, since their state rarely shifts mid-recording. Procedural sedation offers a great deal. An endoscopy patient drifts in and out of consciousness under controlled drugs, with clinicians present and annotating.

This is the same logic the anesthesia literature used to mature measures such as perturbational complexity, and it is the logic behind Hakwan Lau and collaborators’ May 2026 Neuron argument, covered on this site in Lau and the scientific standards for AI consciousness, that the field needs covert measures and better instruments before it can say anything rigorous about non-reporting systems. A public corpus of 1,129 annotated transitions is the kind of substrate that argument asks for. It also connects to the ignition-threshold work this site examined in Baars and the global workspace ignition thresholds, where the empirical grounding of ignition-style measures depends on exactly such state-change data.

Limits

Three limits bound what DOSE-I can support. The population is endoscopy patients, so the drug profiles, ages, and clinical contexts are narrow. The annotations are clinical judgments of sedation depth, made for care rather than for consciousness science, so the labels carry the assumptions of that clinical scale. And sedation is one path into unconsciousness. Findings validated on drug-induced transitions need separate validation for sleep, disorders of consciousness, and anesthesia generally. The dataset’s own framing is technical, and the paper claims a resource, not a theory.

Comparison to The Consciousness AI

The Consciousness AI project’s research stance holds that consciousness is an emergent, substrate-independent property, and its measurement discipline requires that any proposed indicator be checked against known cases before it is pointed at novel systems. DOSE-I widens the supply of known cases. A transition-annotated corpus does not settle anything about machines. It strengthens the ground truth layer that both human and machine consciousness measures will be tested against, which is the shared dependency underneath every framework the project tracks. The project runs no clinical data pipeline, so the dataset supports its research program at the level of method, not of code.

What It Changes and What It Leaves Open

The dataset changes who can work on transition detection, because biosignal sedation data of this size was previously scattered across clinical archives. Benchmark results on DOSE-I will make indicator proposals comparable, since they will share one annotated ground truth. What it leaves open is the distance between sedation depth and consciousness itself, the generalization from one drug regime to the many ways consciousness levels shift, and the translation from clinical annotation to the finer states consciousness science distinguishes. Those are the next datasets’ problems. The first public 1,129 are now on the table. Its place in the wider measurement debate is mapped in [AI Consciousness in 2026, the current state of the field]](/posts/scientists-race-define-ai-consciousness-2026/).

The dataset DOSE-I by Jakob Garbe, Jan W. Kantelhardt, Katja Seeliger, and Thomas Schmid was posted to arXiv on June 30, 2026 as arXiv:2606.02570, with the primary citation at AIME 2026 (Springer) and data plus preprocessing code on Zenodo.

Researchers covered here