This page replays one recorded episode of one training run. You see the frames the agent received, the bids each module made, the winner the workspace picked, the sync signal between modules, and the reward that came back. The agent is the mark at the centre of every frame, two rings and a dot, drawn in the colour of this site. It is drawn into the picture the agent receives, so it is part of the run and not decoration.
In this episode the two senses are close. Hearing wins 0.435 of the ignited steps and vision wins 0.565. The light is in view on 13 of 200 steps and the agent reaches it on 4 steps. The workspace stays silent on 0.380 of steps.
This episode was not picked by eye. Gate B4 ran three seeds, and one episode of each seed is published. For every seed the published episode is the one whose audio share is closest to that seed’s audio share over the whole run. The three replays together show the spread between the seeds. Seed 57 is nearly all vision, seed 59 is an even contest, and seed 58 sits between them.
The run belongs to Gate B4, and the gate PASSED every criterion at all three seeds. Three limits stand with that pass. It is not a repeat of the earlier gate, because the seeds and the pixels differ. The share won by hearing also rises with the episode number at two of the three seeds, so this design does not separate the task link from learning over time. A passed gate is not evidence of feelings.
The view is a window of 48 pixels around the agent in a room of 224 pixels, so the world moves and the agent does not. The colours are the ones this run was drawn with.
Outlined bar, raw bid. Filled bar, bound bid. All five rows share one scale, the highest bid of the episode, so bar lengths can be compared between modules. A raw bid never passes 1.0. A bound bid can pass it, because the affective modulator multiplies the raw bid. The colour marks the module, never a state of the world.
Dark strip cells are steps with no winner. Each line is drawn between its own lowest and highest value, which are printed inside it, so a line that fills the box can still be a flat line over a tiny range. The lower line is the SHAPED reward. That is the environment reward plus the curiosity terms the training loop adds. Tap the strip to move the playhead.
The manifest is written by the ethics framework at run start. The framework holds rules E1 to E8 in the framework document; the rules listed above are the ones a run must satisfy before its first step, and the others are checked by a person. No session on this page is evidence of feelings or awareness. Finding a light is automatic approach, which Feinberg and Mallatt exclude as evidence of affect. The consciousness clock never moves because of a session.
| Environment | dark_room |
|---|---|
| Seed | 59 |
| Episode | 4 of a 10 episode, 200 step run |
| Ethics framework | v1.0, existence drive on |
| Flags | --enable-audio --rssm-latent-mode continuous --capsule-workspace-source all_levels --existence-drive on --dark-room-audio binaural --dark-room-audio-channels 4 --dark-room-collision --dark-room-view agent_centered --vision-bid-reduction zscore --audio-salience surprise --learned-valence --ignition-rule tolerance --ignition-tolerance-sd 1.0 --bid-precision gain --dark-room-agent-mark ring --dark-room-agent-colour 217,119,87 |
| Measured result | PASSED every criterion. Gate B4 ran the same criteria at three seeds, with the agent drawn as the project mark. Competition, silence, selectivity and the kill rule all passed, and so did the task criterion (pooled rho 0.458, one-sided p 0.006). Three limits stand. This is not a repeat of Gate B3, because the seeds and the pixels differ. The share won by hearing still rises with the episode number at two of the three seeds. A passed gate is not evidence of affect. |
| Verdict | verdict document in the public research repository |
The command and the seed above are the whole recipe the run started from. The sound is seeded too. Whether a second run with the same command repeats this episode step by step was not measured for these flags.
The winner strip and the ignition readout are instruments, not reports of experience. A module winning a bid is a circuit event in a trained network. Finding a light is automatic approach, and Feinberg and Mallatt exclude automatic approach and avoidance as evidence of affect. Nothing on this page claims feelings, awareness or consciousness, and no session moves the consciousness clock on the architecture page.
The two columns marked ground truth, in light and light in view, come from
the environment’s own state. The agent never receives them. They show the state
that produced the frame of that step, which is one step earlier than the reward.
On the step the agent enters the light, the reward is already 1.0 while in
light is still false.
Every number the console draws is read at run time from the session’s own
bundle, which holds steps.json, session.json and ethics_manifest.json under
/public/data/training_sessions/dark-room-b4-seed59-episode4/. The bundle was
produced from the recorded run by the repository’s export script, which strips
recording dates and local paths before anything is written. Nothing on this
page was typed by hand.
The other published sessions are listed on the training sessions index.