This page replays one recorded episode of one training run. You see the frames the agent received, the bids each module made, the module with the highest bid at each step, the sync signal between modules, and the reward that came back. The agent is the mark at the centre of every frame, two rings and a dot, drawn in the colour of this site. It is drawn into the picture the agent receives, so it is part of the run and not decoration.
In this episode the light is in view on 6 of 200 steps and the agent never reaches it. Hearing has the highest bid on 0.203 of the ignited steps, 24 of 118, and vision on 0.797. The workspace stays silent on 0.410 of steps. At every one of the 118 ignited steps both modules passed the threshold, and the policy received the vector of the module with the lower bid. This is the seed where hearing had the highest bid least often over the whole run, 0.211 of ignited steps.
This episode was not picked by eye. Gate B7 ran 10 seeds, and three of them are published. The seeds were put in rising order of the share of ignited steps where hearing had the highest bid over the whole run, and the lowest, the middle and the highest seed were taken. That rule was written before the result of the first gate of this size was seen. For each seed the published episode is the one whose audio share is closest to that seed’s audio share over the whole run.
The view is a window of 48 pixels around the agent in a room of 224 pixels, so the world moves and the agent does not. The colours are the ones this run was drawn with.
Outlined bar, raw bid. Filled bar, bound bid. All five rows share one scale, the highest bid of the episode, so bar lengths can be compared between modules. A raw bid never passes 1.0. A bound bid can pass it, because the affective modulator multiplies the raw bid. The colour marks the module, never a state of the world.
Each strip cell shows the module with the highest bid at that step. Dark cells are steps where the workspace is silent. The highest bid does not say which vector the policy received. The session page states that. Each line is drawn between its own lowest and highest value, which are printed inside it, so a line that fills the box can still be a flat line over a tiny range. The lower line is the SHAPED reward. That is the environment reward plus the curiosity terms the training loop adds. Tap or drag along the strip to move the playhead.
The manifest is written by the ethics framework at run start. The framework holds rules E1 to E8 in the framework document; the rules listed above are the ones a run must satisfy before its first step, and the others are checked by a person. No session on this page is evidence of feelings or awareness. Finding a light is automatic approach, which Feinberg and Mallatt exclude as evidence of affect. The consciousness clock never moves because of a session.
Gate B6 failed on competition at 3 of its 10 seeds. At those seeds vision had the highest bid at nearly every step. A test at the same 3 seeds found the cause. One part of the agent, the learned valence, gives each sense a bid boost that grows with reward and has no upper limit. Gate B7 asked the Gate B6 questions again at 10 seeds never run before, with that part switched off. Its two questions and its limits were written before the runs.
The run belongs to Gate B7. The gate PASSED both questions. At all 10 seeds no module had the highest bid at 0.95 or more of the ignited steps. The top share is 0.522 to 0.789. Pooled over 100 episodes, hearing has the highest bid at a larger share of the ignited steps in episodes where the light is out of view (rho 0.388, one sided p 0.0005, rho above 0 at 8 of 10 seeds). It is the first pass of the full gate at 10 seeds.
Four limits stand with that result.
The gate judges which module has the highest bid. It does not judge which vector the workspace sends to the policy. A later measurement on these 10 runs found that two modules pass the threshold at 0.905 to 0.974 of the ignited steps, and that the policy then receives the vector of the module with the lower bid. The policy received the vector of the module with the highest bid at 0.026 to 0.095 of the ignited steps.
The share of hearing also rises with the episode number at all 10 seeds, so this design does not separate the task link from learning over time.
Gate B6 and Gate B7 use different seeds. The pair does not show what the learned valence adds or removes at a given seed.
A passed criterion is not evidence of feelings.
| Environment | dark_room |
|---|---|
| Seed | 81 |
| Episode | 4 of a 10 episode, 200 step run |
| Ethics framework | v1.0, existence drive on |
| Flags | --enable-audio --rssm-latent-mode continuous --capsule-workspace-source all_levels --existence-drive on --dark-room-audio binaural --dark-room-audio-channels 4 --dark-room-collision --dark-room-view agent_centered --vision-bid-reduction zscore --audio-salience surprise --ignition-rule tolerance --ignition-tolerance-sd 1.0 --bid-precision gain --dark-room-agent-mark ring --dark-room-agent-colour 217,119,87 |
| Measured result | PASSED both questions at 10 seeds. Gate B7 asked the Gate B6 questions at 10 fresh seeds, without the learned valence. Both were fixed before the runs. Competition at every seed PASSED (the top share is 0.522 to 0.789 and the kill rule fired at no seed). The task link PASSED (pooled rho 0.388 over 100 episodes, one-sided p 0.0005, rho above 0 at 8 of 10 seeds). Four limits stand. The gate judges which module has the highest bid, and the policy received the vector of that module at 0.026 to 0.095 of the ignited steps. The share of hearing also rises with the episode number at all 10 seeds. Gate B6 and Gate B7 use different seeds. A passed criterion is not evidence of affect. |
| Verdict | verdict document in the public research repository |
The command and the seed above are the whole recipe the run started from. The sound is seeded too. Two runs with the same command and seed gave identical files in a check at another seed. That was not checked at this seed.
The highest bid strip and the ignition readout are instruments, not reports of experience. A module having the highest bid is a circuit event in a trained network. Finding a light is automatic approach, and Feinberg and Mallatt exclude automatic approach and avoidance as evidence of affect. Nothing on this page claims feelings, awareness or consciousness, and no session moves the consciousness clock on the architecture page.
The two columns marked ground truth, in light and light in view, come from
the environment’s own state. The agent never receives them. They show the state
that produced the frame of that step, which is one step earlier than the reward.
On the step the agent enters the light, the reward is already 1.0 while in
light is still false.
Every number the console draws is read at run time from the session’s own
bundle, which holds steps.json, session.json and ethics_manifest.json under
/public/data/training_sessions/dark-room-b7-seed81-episode4/. The bundle was
produced from the recorded run by the repository’s export script, which strips
recording dates and local paths before anything is written. Nothing on this
page was typed by hand.
The other published sessions are listed on the training sessions index.