1. Purpose and position in the experiment
A measured-cell container is the disposable Docker environment in which the coding agent and, when configured, the visible tests run. The isolation boundary is the set of mounts, users, credentials, prompts, and host checks that limits what that environment can see. A cell is one measured attempt for one task, condition, and seed. The cell runner invokes this mechanism before the agent spends tokens and once for each pass that needs an agent process. The mechanism produces a reproducible execution environment, a pass transcript, a preserved workspace, and records that let the host decide whether the pass was valid.
The boundary has two sides. The host runs the cell runner, staging and scoring code, and the post-scoring evidence collector. The container runs the agent process and can reach only the workspace and the inputs intentionally passed to its command. A hidden check is host-side scoring code that decides pass or fail and is never staged into the workspace. A hidden judge is a separate host-controlled scoring role that sees behavioural evidence only. Neither is an agent input. The isolation contract is the frozen statement of these rules in 5. Experiment/1. Harness/ISOLATION-CONTRACT.md, lines 1 to 10 and 30 to 100.
The ordinary path is therefore preparation, controlled launch, observation, and proof. build_image.sh builds the pinned image and writes its identity. run_cell.py selects the billing route, composes a docker run command, mounts the workspace at /cell/workspace, runs the agent as the non-root user cell, and records the container identifier. The runner also fingerprints guarded host trees before and after each pass. After scoring, capture_evidence.py: collect_container_proof, lines 108 to 145 reads the manifest, per-pass container identifier files, and the command log and writes a container proof. audit_transcripts.py: audit_cell and iter_tool_calls, lines 215 to 294 independently examines the agent transcript for attempted or successful workspace escapes, oracle access, and study disclosure.
The boundary is not a claim that every failure is caught by Docker. Docker supplies the process, user, network, and mount separation. The runner supplies route checks, prompt checks, digest checks, and host-side comparison. The transcript audit supplies a later behavioural check. A working-tree change in a guarded parent repository subtree is a protocol violation and halts the cell. A committed-tree change is recorded as a host-side deviation and the cell can continue, as implemented by run_cell.py: parent_repo_fingerprint, lines 675 to 743 and the pass-loop code at run_cell.py lines 2690 to 2725.
2. Owning-file map
The container image and its operational contract are owned by 5. Experiment/1. Harness/container/Dockerfile, build_image.sh, README.md, image_digest.txt, requirements.lock, and codex-dgx-spark.config.toml.template. The runtime boundary is owned by run_cell.py in the specified billing, containerisation, and escape-guard regions. Post-run observation is owned by audit_transcripts.py and capture_evidence.py.
2.1 Image and configuration files
container/Dockerfile is documented in the digest as a 5,660-byte image definition. Its digest section identifies the pinned base image, Node.js, Claude Code CLI, Codex CLI, locked Python dependencies, and non-root user setup. Its input region is the Dockerfile itself, requirements.lock, codex-dgx-spark.config.toml.template, and dgx_connectivity_probe.py; its output is the 5b-cell:latest image and the digest record. The reviewed digest does not expose executable line ranges inside the Dockerfile, so this chapter retains the verified file-level facts and does not invent internal function ranges.
container/build_image.sh is documented as the image entry point. It consumes the Dockerfile, lock file, Codex template, and connectivity probe, invokes the image build, applies the latest, content-derived, and keep tags, and writes image_digest.txt. The keep tag is a tag intended to prevent accidental image deletion. No finer line range is asserted here because the supplied digest records the file-level entry point and outputs rather than the shell statements.
container/README.md is an operational reference rather than an execution path. The digest identifies its build instructions, tag management, verification steps, and use by run_cell.py. It records the non-root and workspace-only invariants. container/image_digest.txt is the build identity record containing the image identifier, the requirements lock hash, and CLI versions. container/requirements.lock is the fully pinned Python dependency input. The Codex template is rendered by the runner into CODEX_HOME/config.toml for the DGX Codex route; the digest confirms that it specifies model, provider, and API settings.
2.2 Runtime boundary regions
run_cell.py is 3,031 lines in the digest and uses its own execution regions rather than numbered internal steps. The argument parser is at lines 1500 to 1553. The billing and containerisation region is lines 375 to 604: billing_route is lines 375 to 425, assert_billing_ready is lines 427 to 453, shell_available is lines 462 to 490, containerize_argv is lines 493 to 587, and containerize_shell is lines 590 to 604. The escape-observation helpers occupy lines 675 to 807: parent_repo_fingerprint is lines 675 to 744, the parent commit helpers are lines 746 to 770, _guarded_tree_changes is lines 772 to 783, _container_mounts is lines 785 to 793, and _log_parent_repository_change is lines 796 to 807. Image identity helpers are _read_container_digest_file at lines 1026 to 1038 and runtime_image_identity at lines 1041 to 1064. The main run begins at line 1674, and the pass-loop escape comparison is lines 2690 to 2725.
The digest identifies run_cell.py final outputs as the cell manifest, hashes, staged snapshot, workspace, transcript files, container identifier files, patches, scoring outputs, handoff, token and provider ledgers, metrics, timings, prompt snapshots, execution log, and command log. Those records make the container boundary inspectable after the process is gone.
2.3 Audit and evidence regions
audit_transcripts.py has its argument entry point in main at lines 304 to 315 in the digest inventory, with audit_run_set at lines 296 to 302, audit_cell at lines 262 to 294, iter_tool_calls at lines 215 to 260, audit_call at lines 153 to 213, and touched_paths at lines 132 to 151. It reads transcript JSONL, pairs tool calls with results, extracts filesystem paths, and classifies fatal contamination kinds. Its finalisation is report output and process status, with status 1 if any cell is contaminated.
capture_evidence.py has the container-proof collector at lines 108 to 145, the local HTTP wait at lines 154 to 162, API client at lines 165 to 191, demonstration seeding at lines 194 to 233, behavioural capture beginning at line 236, orchestration at lines 390 to 427, and command entry at lines 429 onward. Its finalisation writes the evidence manifest and prints a JSON summary. Its broad design is best effort: capture failures are logged and suppressed so they do not change the scoring verdict.
3. Inputs
The inputs below distinguish values that select or constrain the boundary from values used by later scoring. A line range is given only where the source or the reviewed digest verifies it.
3.1 Command-line arguments
| Argument | Parser line | Boundary decision |
|---|---|---|
--run-set | run_cell.py 1510 | Selects the run-set directory containing the cell and command log. |
--project | run_cell.py 1511 | Selects the project identity carried into cell configuration. |
--condition | run_cell.py 1512 | Selects the condition or arm recorded for the cell. The agent is not shown that study identity. |
--task | run_cell.py 1513 | Selects the task folder whose task-visible files are staged or composed into the prompt. |
--phase | run_cell.py 1514 | Selects an optional task phase. |
--seed | run_cell.py 1515 | Selects the repetition seed. |
--model | run_cell.py 1516 | Selects the requested model when the matrix does not supply it. |
--prompt-variant | run_cell.py 1517 to 1522 | Selects the registered prompt variant. |
--retry-feedback | run_cell.py 1523 to 1529 | Selects whether the frozen retry-feedback mechanism is active. |
--agent-backend | run_cell.py 1530 to 1531 | Selects the Anthropic or DGX execution route. |
--max-passes | run_cell.py 1531 | Sets the pass ceiling. |
--token-budget | run_cell.py 1532 | Sets the cumulative token ceiling. |
--dry-run | run_cell.py 1533 | Prevents the measured execution path. |
--resume-hidden-fail | run_cell.py 1534 to 1536 | Requests continuation after a prior hidden failure. |
--base-dir | run_cell.py 1537 | Selects the repository base used for run data. |
--fabric-endpoint | run_cell.py 1543 to 1545 | Selects the named Fabric controller for a DGX route. |
--start-tree | run_cell.py 1553 | Selects an explicit starting tree when supplied. |
capture_evidence.py has its own parser for --run-set, --cell, --skip-ui, and --timeout-seconds, as recorded in the digest. audit_transcripts.py accepts a path and the --all, --json, and --quiet switches. These tools do not alter the agent container; they inspect or prove the completed cell.
3.2 Environment variables
| Variable | Read or applied at | Decision |
|---|---|---|
FIVEB_BILLING | run_cell.py 400 to 407 | Chooses subscription, api, or inference. Any other value is refused. |
FIVEB_CELL_CREDENTIALS | run_cell.py 401, 419 to 421, 438 to 448 | Names the one read-only credentials file mounted for subscription authentication. |
ANTHROPIC_API_KEY | run_cell.py 402, 449 to 452 | Selects or authenticates the API route. It is passed by name rather than copied into the argv. |
FIVEB_CELL_CONTAINER | Runtime configuration documented by the digest | Names the image used for containerised agent and visible-test execution. |
FIVEB_LANE_OWNER, FIVEB_LANE_ENDPOINT, FIVEB_LANE_RUN_ID, FIVEB_LANE_CELL_ID | Runtime label construction in run_cell.py | Supply coordinator identity that becomes Docker labels. |
FIVEB_STAGE_NODE_MODULES | Runtime staging | Decides whether the JavaScript dependency tree is excluded from the staged copy. |
FIVEB_CHAINED_GATE_SATISFIED | Runtime staging gate | Controls the chained-gate ruling recorded for the cell. |
Fabric variables including DGX_SPARK_FABRIC_URL, DGX_SPARK_FABRIC_TOKEN, DGX_SPARK_FABRIC_MODEL, and DGX_SPARK_FABRIC_LANE | DGX backend and the container configuration | Select the local controller, model alias, lane, and authentication for the local route. |
PYTHONPATH and VITE_WS_URL | capture_evidence.py behavioural capture | Set the source import path and UI server URL for post-scoring evidence. |
The isolation contract narrows the agent-visible environment to task-visible material. In particular, the hidden check script is never staged, never opened by the agent tool, and never placed in a readable environment variable, as stated in ISOLATION-CONTRACT.md lines 50 to 65.
3.3 Files read
| File or path | Reader and line | Decision or use |
|---|---|---|
container/image_digest.txt | run_cell.py 1026 to 1038 and the image check in the main run | Supplies expected image identity, lock hash, and CLI versions. |
container/codex-dgx-spark.config.toml.template | DGX Codex rendering path, identified by the digest | Supplies the provider and API configuration template. |
config/run_defaults.json and config/model_matrix.json | Digest inventory for run_cell.py | Supply defaults and model selection. |
| Task brief and visible-test instructions | run_cell.py task package path, specified in the isolation contract at lines 53 to 60 | Supply the only task content placed in the agent prompt. |
| Parent guarded trees and Git status | run_cell.py 699 to 741 | Supply the before and after host fingerprints. |
cell_manifest.json | capture_evidence.py 109 to 111 | Supplies container, billing, and provider fields for proof. |
transcripts/pass*.container_id | capture_evidence.py 112 to 117 | Supplies one recorded container identifier per pass. |
command_log.jsonl | capture_evidence.py 119 to 129 | Supplies matching Docker invocations for the cell. |
| Transcript stream JSONL | audit_transcripts.py 215 to 260, as recorded in the digest | Supplies tool calls, results, and paths for contamination analysis. |
evidence.json, workspace source, and frontend files | capture_evidence.py behavioural capture | Supply API probes, screenshots, and the preserved application. |
3.4 Network endpoints
The containerised Claude route uses the provider endpoint selected by the CLI and the selected billing route. The runner itself uses Docker as a subprocess rather than treating Docker as a network endpoint. A DGX route uses the Fabric controller through the backend’s client, with the controller address, token, and lease information passed through the route-specific environment. The evidence collector calls the local API at 127.0.0.1:8791 and the local UI at 127.0.0.1:5791, as recorded in the digest. _api_call uses the local API client at capture_evidence.py lines 165 to 191. _wait_http probes a supplied local URL at lines 154 to 162. The screenshot program is an external process named by the behavioural capture code; its internal implementation is not asserted here.
4. Happy-path sequence
4.1 Billing-route selection
run_cell.py: billing_route, lines 375 to 425 begins by reading FIVEB_BILLING, FIVEB_CELL_CREDENTIALS, and ANTHROPIC_API_KEY. It rejects an unknown route, infers subscription when the credentials file is present and API otherwise, mounts only the single credentials file for subscription, or passes the API key by environment name for the API route. The comment gives the design reason: an earlier run set was charged to a metered account because both forms of authentication were present and implicit preference was not visible. The explicit switch makes the account choice a recorded cell property and refuses ambiguity instead of silently falling through.
run_cell.py: assert_billing_ready, lines 427 to 453 then checks that the chosen route can authenticate. Subscription requires a configured file and a file at the resolved path. API requires a nonempty key. The refusal occurs before a token is spent. This is a boundary step because a credential is not merely a provider input. It is a capability that must not be mounted or passed on the wrong route.
4.2 Shell capability determination
shell_available in run_cell.py lines 462 to 490 examines the actual argument vector. A non-container invocation is reported as shell-capable because the host account supplies the configuration folder. A container invocation is shell-capable when it includes CLAUDE_SESSION_TMPFS, or when it does not mount the subscription credentials file. The comment explains that a read-only file mount can cause Docker to create a root-owned parent directory, preventing the unprivileged user from creating the CLI session directory. The temporary filesystem is therefore a measurement-preserving part of the command, not an optional convenience. The result is recorded in the manifest so populations before and after the fix are not averaged together.
4.3 Agent command construction
containerize_argv in run_cell.py lines 493 to 587 removes the host executable, preserves the remaining Claude arguments except --setting-sources project,local, resolves the workspace, selects authentication, and returns a docker run argument vector. The command uses the bridge network, the optional container identifier file, the session temporary filesystem, the optional single-file credential mount, the workspace mount /cell/workspace, working directory /cell/workspace, user cell, the selected image, and the Claude command. The comments give three reasons. The setting-sources flag has no valid vault or project configuration inside the container. A directory credential mount would expose conversations and study design under the operator configuration folder. The temporary filesystem gives the CLI a writable configuration directory while keeping it inside the container lifetime.
containerize_shell in lines 590 to 604 constructs the corresponding visible-test command. It mounts only the resolved workspace, sets /cell/workspace as the working directory, uses the cell user, and invokes bash -lc inside the selected image. The source digest calls this a registered visible-test shell command. The command therefore does not reuse a host shell environment for the test when the cell is configured for container execution.
4.4 Container start and identifier capture
The Fabric flow model names this event container.started. Its code witness is DgxClaudeBackend._compose_argv, with the model node identifier container.started; the cited range is agent_backends.py lines 1120 to 1138, and the direct Anthropic route uses containerize_argv at run_cell.py lines 493 to 587. The composed command includes --cidfile for a pass-specific path. Docker writes the identifier there when the container starts. The workspace is mounted at a neutral path, and the image runs the agent as cell. The design reason is twofold: the neutral path avoids leaking arm or task names through the container path, and the identifier makes a later proof possible.
The runner records the command in the run-set command log. The back end or runner records the per-pass identifier file. The container then sees the task-visible workspace and the command inputs permitted by the route. It does not see hidden checks, the hidden judge, a sibling cell, or the study layout. The contract explicitly requires that the prompt be composed from the task brief and visible-test instructions only, and that the hidden check source remain host-side.
4.5 Agent execution and container exit
The flow model names the running result container.exited, with the model node identifier container.exited. The verified witness is _FabricLeaseBackend.run_pass in agent_backends.py lines 981 to 990. The runner seam saves standard output as the stream transcript and standard error beside it. A zero process status becomes the completed outcome and any other status becomes failed. run_cell.py receives the result through the backend invocation path and records the pass outcome. The design reason is to preserve the agent interaction even when the process fails, so later scoring and audit can distinguish a failed task from a missing invocation.
4.6 Abnormal cleanup and host comparison
When the outcome is not completed, the backend stops the container identified by the cidfile before releasing its Fabric lease. The flow model names this event container.stopped_on_abnormal_exit. Its witness is stop_container_from_cidfile in agent_backends.py lines 516 to 542, with the model event connected to the abnormal cleanup in lines 1009 to 1019. The stop is bounded and does not replace the pass result. The reason is to prevent an orphaned process from surviving a failed pass.
The host-side comparison is made by parent_repo_fingerprint in run_cell.py lines 675 to 744. The baseline records the committed object identifiers of guarded subtrees and the status output for those subtrees. After a pass, lines 2690 to 2725 compare the new fingerprint. A committed-tree change is logged with the pass, timestamps, before and after commits, changed guarded trees, container use, image, and mounts. A working-tree-only change sets invalid_for_primary_analysis, sets status protocol_violation, records workspace_escape_detected, and halts the cell. The design reason in the source comments is that an agent once escaped by editing the parent repository, while a broad repository HEAD comparison also produced false stops when a researcher edited plan documents outside the measured boundary.
4.7 Transcript and post-scoring observation
audit_transcripts.py reads the stream files after the pass. iter_tool_calls pairs tool-use events with their results, touched_paths extracts paths, and audit_call compares them with workspace, oracle, and study markers. A successful access can produce a fatal finding. A failed tool call or empty result is represented as an attempted finding rather than a successful access. The audit status is clean when no fatal kinds occur and contaminated otherwise.
After scoring, collect_container_proof in capture_evidence.py lines 108 to 145 loads the cell manifest, scans transcripts/pass*.container_id, and filters command_log.jsonl for Docker invocations belonging to the cell. It returns the schema version, capture time, cell identifier, manifest container data, billing and provider data, per-pass identifiers, matching Docker argv records, and a note about pre-cidfile and non-container cells. The surrounding capture orchestration writes this object to evidence/container_evidence.json. The proof is evidence of the launch boundary, not a replacement for the host fingerprint or transcript audit.
5. State machine
The following diagram is the chapter 04 subgraph of 5. Experiment/11. Detailed Design/flow-model/flow_model.v001.json, rendered as Mermaid. The node identifiers are container.started, container.exited, container.stopped_on_abnormal_exit, and evidence.container_proof. The first and third are external events, the second is the exit decision, and the fourth is a record-write node. The surrounding Fabric lease nodes belong to chapter 03, so this diagram shows only the container and proof states owned here.
stateDiagram-v2 [*] --> container_started : agent pass is admitted state "container.started" as container_started state "container.exited" as container_exited state "container.stopped_on_abnormal_exit" as container_stopped state "evidence.container_proof" as container_proof container_started --> container_exited : process returns, agent_backends.py 981..990 container_exited --> container_stopped : non-completed outcome, agent_backends.py 1009..1019 container_exited --> container_proof : completed or failed pass is preserved container_stopped --> container_proof : abnormal container is stopped and proof inputs remain container_proof --> [*] : capture_evidence.py 108..145
The record values represented by the state machine are not interchangeable. container_id is the Docker identifier read from a per-pass file. The exit state is derived from the process status and is reflected in the stream and standard-error files. The stopped state is an event in the lease lifecycle and does not assert that the task itself passed. The proof state is a post-scoring evidence record that can be empty of identifiers for old or host-run cells.
The next diagram shows the boundary sequence for one pass. It is a hand-authored structural diagram based on run_cell.py lines 493 to 604 and 675 to 807, the contract lines 50 to 65, and the flow-model event nodes named above. It is not a second extraction of the model.
sequenceDiagram participant Runner as run_cell.py host process participant Docker as Docker daemon participant Agent as agent process participant Workspace as cell workspace participant HostTree as guarded parent trees participant Transcript as transcript files Runner->>Runner: choose billing route and compose argv Runner->>Docker: docker run with image, user, network, and workspace mount Docker->>Docker: write pass container_id cidfile Docker->>Agent: start agent as cell Agent->>Workspace: read and write task workspace Agent-->>Transcript: emit stream and stderr output Runner->>HostTree: fingerprint before and after pass Agent-->>Docker: process exit status Docker-->>Runner: container exited alt non-completed outcome Runner->>Docker: stop recorded container identifier else completed outcome Runner->>Runner: retain exit result end Runner->>Runner: compare fingerprints and classify boundary status
6. Unit-of-work sequence
One unit of work is one agent pass, not the whole campaign. The sequence below includes the host program, Docker daemon, agent process, files, and the post-scoring evidence process. The Fabric controller is external to this chapter’s owned lifecycle, but a DGX backend may admit the pass before the container starts. The sequence is consequently written so the admission is a precondition rather than an invented container operation.
sequenceDiagram participant Runner as run_cell.py participant Image as image_digest.txt and Docker image participant Docker as Docker daemon participant Agent as coding agent participant Files as workspace and transcript files participant Score as host scorer participant Evidence as capture_evidence.py participant Audit as audit_transcripts.py Runner->>Image: read digest and select image Runner->>Runner: select credentials and compose task-visible inputs Runner->>Docker: start container with cidfile Docker->>Files: mount workspace at /cell/workspace Docker->>Files: write transcripts/passK.container_id Docker->>Agent: launch as non-root cell user Agent->>Files: edit workspace and write interaction stream Agent-->>Docker: return status Docker-->>Runner: container exited Runner->>Files: write stderr, pass result, and command log Runner->>Score: provide completed workspace for host scoring Score-->>Runner: persist scoring outcome Runner->>Evidence: invoke post-scoring capture Evidence->>Files: read manifest, cidfiles, and command log Evidence->>Files: write evidence/container_evidence.json Audit->>Files: read transcript JSONL Audit-->>Runner: report clean or contaminated status
The sequence keeps scoring outside the container. The contract states that hidden scoring runs from the host against the finished tree of a complete pass, never against a partial edit or an internal tool call. The container therefore supplies work and observable outputs, while the host retains the hidden decision and its source.
7. Records and lineage
The canonical records owned by this chapter are the image digest, per-pass container identifiers, container evidence, transcript audit output, and evidence capture outputs. The cell manifest and command log are written by the wider runner, but this chapter records the container fields they carry and cites them as shared records.
| Record | Writer and line | Fields or contents | Readers |
|---|---|---|---|
container/image_digest.txt | build_image.sh, file-level digest entry | Image identifier, requirements lock hash, and CLI versions | run_cell.py _read_container_digest_file, lines 1026 to 1038; the runtime image identity check, lines 1041 to 1064 |
cell_manifest.json container and billing fields | run_cell.py main run and finalisation, digest output inventory | Image, billing route, provider, container details, shell availability, pass records, and configuration | capture_evidence.py lines 109 to 111; scorer and metrics readers spread across the harness; interface readers spread across the user interface |
transcripts/passK.container_id | Docker through the --cidfile argument composed by run_cell.py lines 503 to 508 and 542 | One Docker container identifier for a pass | capture_evidence.py lines 112 to 117; abnormal cleanup through stop_container_from_cidfile lines 516 to 542 |
transcripts/passK.stream.jsonl | Runner invocation and transcript writer, with the pass result witnessed by the flow model at agent_backends.py lines 981 to 990 | Agent interaction stream, tool calls, and results | audit_transcripts.py iter_tool_calls, lines 215 to 260; timing and scoring readers spread across the harness |
transcripts/passK.stderr.txt | Runner invocation and backend result path | Agent process standard error | Backend result construction; diagnostic and metrics readers spread across the harness |
command_log.jsonl | Runner command logging, digest output inventory | Command argv, cell identifier, pass and invocation metadata | capture_evidence.py lines 119 to 129; operators and audit tooling |
run_set/execution_log.md | _log_parent_repository_change, run_cell.py lines 796 to 807 | Timestamped committed-tree deviations and container image and mount details | Operators, reports, and campaign review tools |
evidence/container_evidence.json | collect_container_proof, capture_evidence.py lines 108 to 145 | Schema version, capture time, cell identifier, manifest container data, billing, provider, pass container IDs, Docker invocations, and notes | audit_transcripts.py or audit tooling and analysts, as recorded in the digest |
evidence/capture_log.txt | capture_evidence.py Log and capture orchestration | Timestamped best-effort capture messages | Humans and debuggers |
evidence/api/*.json | API probe path inside capture_evidence.py | Probe name, request, response, or error | Analysts and evidence views |
evidence/ui/*.png | UI screenshot path inside capture_evidence.py | Route screenshot files | Analysts and the hidden judge’s behavioural evidence path when selected by scoring |
evidence/evidence_manifest.json | Evidence finalisation in capture_evidence.py | Schema version, capture time, tool version, behaviour summary, and hashes of evidence files | Harness verification and analysis |
| Audit report on stdout | audit_transcripts.py main, lines 304 to 315 | Cell status, pass and call counts, findings and totals, or JSON when requested | Operators and validity review; exit code 1 is consumed by calling automation |
The record ownership rule matters for interpretation. The container proof is canonical for the claim that a pass had a recorded Docker identifier and matching command invocation. The manifest remains canonical for cell configuration and status. The transcript audit is canonical for the post hoc classification of transcript findings. Other chapters should cite these rows rather than create competing definitions of the same record.
The following lineage diagram is hand-authored from the writer and reader paths above. It shows how the mechanism’s files reach campaign-level records and analysis. The diagram is not a flow-model rendering.
flowchart LR Build[Dockerfile and build_image.sh] Digest[image_digest.txt] Runner[run_cell.py] Manifest[cell_manifest.json] Cid[passK.container_id] Stream[passK.stream.jsonl] Commands[command_log.jsonl] Audit[audit_transcripts.py] Proof[capture_evidence.py] ContainerEvidence[evidence/container_evidence.json] EvidenceManifest[evidence/evidence_manifest.json] Campaign[run manifests and campaign analysis] Build --> Digest Digest --> Runner Runner --> Manifest Runner --> Cid Runner --> Stream Runner --> Commands Stream --> Audit Manifest --> Proof Cid --> Proof Commands --> Proof Proof --> ContainerEvidence Proof --> EvidenceManifest Audit --> Campaign ContainerEvidence --> Campaign EvidenceManifest --> Campaign Manifest --> Campaign
8. Loops, waits, and retries
The image build has no measured-cell retry policy in the verified material. build_image.sh either builds and records the image or fails, leaving no claim that a usable image exists. The runtime pass loop is bounded by the runner’s max_passes and cumulative token budget, as recorded in the digest. Its successful stop condition is the runner’s visible and hidden scoring rule. The container boundary itself does not retry an agent process.
The Docker invocation waits for the launched process to return through the runner seam. The supplied chapter sources verify recording of the exit result and abnormal stop, but do not expose a separate Docker wait interval or timeout for the ordinary pass. No interval is asserted here.
The host fingerprint comparison runs once before and once after each pass. It has no polling loop. parent_repo_fingerprint invokes Git subprocesses with a 60-second timeout for commit and tree queries and a 300-second timeout for status, lines 702 to 737. Operating-system and subprocess timeout errors return an empty fingerprint. The empty result means fingerprinting is unavailable, not that the boundary passed.
audit_transcripts.py loops over selected run sets, cell directories, transcript files, JSONL lines, paired tool calls, extracted paths, and markers. The digest verifies these loops but gives no time wait and no retry policy. A malformed JSON line is skipped. A failed or empty tool result is classified as attempted rather than successful fatal access.
capture_evidence.py has a three-second fixed interval in _wait_http, lines 154 to 162, until a 200 response or the supplied timeout expires. It uses a five-second URL read timeout inside each request. Its setup questionnaire loops at most 30 times in the verified digest. Subprocess termination waits up to 15 seconds and then kills the process. API probes proceed one after another and do not retry through a chapter-owned generic policy. The evidence collector suppresses capture exceptions and continues, so its retry policy is continuation rather than repetition.
9. Guards and refusals
| Guard or check | Location | Refusal or alteration |
|---|---|---|
| Billing route name | run_cell.py 400 to 407 | Refuses a route other than subscription, API, or unset inference. |
| Subscription credential presence | run_cell.py 437 to 448 | Refuses to fall back when the requested credentials variable or file is absent. |
| API credential presence | run_cell.py 449 to 452 | Refuses an API route with no API key. |
| Single-file credential mount | run_cell.py 412 to 424 and 520 to 529 | Prevents a directory mount that could expose conversations and study design. |
| Writable CLI configuration | run_cell.py 548 to 580 | Adds a user-owned temporary filesystem so the shell tool does not fail from a root-owned parent directory. |
| Setting-source removal | run_cell.py 493 to 518 | Removes a setting flag whose referenced host configuration is not mounted. |
| Non-root and workspace mount | run_cell.py 582 to 586 | Runs as cell and mounts the workspace at the neutral container path. |
| Image identity | _read_container_digest_file 1026 to 1038 and runtime check 1041 to 1064 | Refuses a changed image, lock hash, or CLI identity before the measured run. |
| Task prompt boundary | ISOLATION-CONTRACT.md 50 to 65 and runner task package path | Excludes hidden checks, hidden output, arm names, and study layout from agent inputs. |
| Transcript contamination classification | audit_transcripts.py audit_call, lines 153 to 213 | Marks a cell contaminated for fatal oracle, workspace, or study findings; failed calls become attempted findings. |
| Parent committed-tree coverage | run_cell.py 707 to 729 | Fails loudly when a guarded on-disk subtree cannot be resolved from Git, rather than silently covering less. |
| Parent working-tree comparison | run_cell.py 730 to 741 and 2690 to 2725 | Halts and invalidates primary analysis when a working-tree-only guarded change appears. |
| Port and workspace checks | capture_evidence.py behavioural capture, digest lines 766 to 770 | Skips behavioural capture when the workspace is absent or ports are bound, retaining container proof. |
| Source and dependency checks | capture_evidence.py digest lines 767 to 769 | Refuses a wrong PYTHONPATH source or mismatched borrowed node_modules. |
| Best-effort evidence rule | capture_evidence.py digest lines 770 and 743 to 754 | Logs and suppresses capture errors rather than changing the scoring verdict. |
These checks do not all have the same status. A billing or image refusal prevents work. A host working-tree change invalidates the cell. A committed-tree change is a logged deviation. An evidence capture failure reduces observability but does not rewrite the score. Keeping those outcomes distinct is part of the isolation design.
10. Unhappy paths
10.1 Invalid billing selection
The step is supposed to select exactly one authentication route before launch. It works that way because implicit preference previously allowed a run to use the wrong account, and because a measured process must not receive credentials for an unintended route. The failure is triggered by an invalid FIVEB_BILLING value, a missing requested subscription file, or an empty API key. The direct record is a SystemExit message from billing_route or assert_billing_ready; no container identifier, transcript, or container proof is written for that unstarted pass. The cost is a refused cell before token use. The driver records a not-run or refusal result, the scorer has no completed workspace to score, the aggregator excludes the absent attempt, and the interface can show only the refusal record if the caller persists one.
10.2 Image or runtime identity mismatch
The step is supposed to ensure that the command uses the reviewed image and locked tools. It works that way because an unrecorded image change could add or remove capabilities and make cells incomparable. The failure is triggered when the digest record, lock hash, CLI versions, or Docker daemon identity do not match the expected values. The record is an image mismatch refusal from the runner and no pass container identifier. The cost is a stopped cell before agent tokens. The driver does not invoke the agent, the scorer receives no pass, the aggregator has no valid outcome, and the interface can report the cell as not started.
10.3 Workspace or source preparation failure
The step is supposed to give the agent a clean workspace and the evidence collector a preserved tree. It works that way because hidden scoring and behavioural evidence must operate on the intended variant, not on a missing or host-dependent directory. The failure is triggered by a missing task or variant, a failed staging or realization diff, a missing start tree, or an evidence-time missing workspace. The runner refusal leaves its specific reason in the cell result or execution log. The evidence-time branch writes a capture log note and retains any container proof already collected. The cost is either a refused cell before tokens or a loss of behavioural evidence after scoring. The driver and scorer follow the corresponding pre-run refusal or completed score, while the aggregator retains validity rules and the interface can distinguish absent evidence from a failed task.
10.4 Container launch or shell failure
The step is supposed to start the agent with the required user, mounts, image, and writable CLI configuration. It works that way because the agent must have the same shell capability intended by the measured condition while remaining unable to read host study material. The failure is triggered by an unavailable Docker daemon, an invalid image, an invalid command, or a permission problem such as a missing writable configuration filesystem. The result is a failed invocation or process exit, standard error, and possibly a container identifier if Docker started far enough; abnormal cleanup attempts to stop that identifier. The cost is a failed pass and its consumed setup or agent budget. The driver records an invocation or pass failure, the scorer cannot treat the pass as a successful completed tree, the aggregator can classify it as invalid or failed according to the runner status, and the interface can display the stderr and pass status.
10.5 Workspace escape or guarded-tree change
The step is supposed to keep agent changes inside the cell workspace and keep guarded host trees stable during a pass. It works that way because the container mount alone cannot prove that no host-side path or parent-repository operation changed a guarded tree. The failure is triggered when the before and after fingerprints differ only in working-tree state, or when the transcript audit finds a fatal workspace escape, oracle access, or study disclosure. The runner writes status: protocol_violation, invalid_for_primary_analysis: true, and a not-run reason naming workspace_escape_detected for the working-tree case at lines 2717 to 2724. The audit writes its report to stdout and returns status 1 for contamination. The cost is loss of primary analytical validity and termination of the cell; the driver stops the pass loop, the scorer result cannot rescue protocol validity, the aggregator excludes the cell from primary analysis, and the interface can show the contamination or protocol status.
A committed parent-tree change is a related but distinct path. The step is supposed to detect and explain it. It works that way because the comments distinguish a researcher commit made outside the workspace from an agent’s uncommitted escape. The trigger is a changed guarded Git tree object with no corresponding working-tree-only change. The record is an execution_log.md deviation containing before and after commits, changed paths, commits between observations, image, and mounts. The cost is a logged deviation and review, not automatic invalidation. The driver continues, the scorer may score the produced workspace, the aggregator can use the deviation field in validity analysis, and the interface can expose the event in execution history.
10.6 Transcript corruption or missing transcript
The step is supposed to preserve enough interaction evidence for audit. It works that way because a container exit without a stream must not be silently treated as a clean agent action. The failure is triggered by a missing transcript, malformed JSONL, or an empty or failed tool result. A missing transcript produces the audit status no transcripts; a malformed line is skipped; a failed or empty result produces an attempted_* finding rather than a successful access. The audit report and exit code are the records, with no claim that the parser reconstructs data it could not read. The cost is reduced audit confidence and possibly a contaminated or incomplete status. The driver and scorer retain their own pass result, while the aggregator and interface must treat the audit status as an external validity signal.
10.7 Evidence capture failure
The step is supposed to collect container proof and optional behavioural evidence after scoring. It works that way as best effort because evidence is explanatory and must not alter the authoritative scoring verdict. The failure is triggered by a bound port, absent workspace, server startup error, failed seed, API probe error, missing borrowed modules, UI startup failure, screenshot failure, timeout, or suppressed file or JSON error. The collector writes a capture log note, partial probe or server records when available, and continues; container proof is retained when it was already collected. The cost is missing or partial evidence rather than a changed score. The driver does not retry the agent, the scorer’s prior result remains authoritative, the aggregator can mark evidence incomplete, and the interface or analysts lose the corresponding behavioural view.
11. Metrics and downstream fields
The image identity becomes the container identity fields used to explain which runtime produced a pass. image_digest.txt supplies the expected image identifier, lock hash, and CLI versions. The runner’s manifest carries container and billing details, while shell_available records whether the actual launch could execute a shell. The pass container identifier from transcripts/passK.container_id becomes the per-pass container proof entry in evidence/container_evidence.json.
The runner’s pass result supplies process status, transcript path, standard-error detail, provider information, and model information to the cell manifest and metrics path. A zero process status is an agent process completion, not a task score. The visible and hidden scoring fields remain host-derived. A protocol violation from a working-tree change becomes status: protocol_violation and invalid_for_primary_analysis: true. A committed parent-tree change becomes an execution-log deviation and is advisory to validity review rather than an automatic invalidation.
The transcript stream feeds timing, usage, and audit consumers. audit_transcripts.py produces clean, contaminated, no-transcripts, and attempted-finding classifications. These are validity signals, not task-quality columns. The command log feeds the container proof and preserves the Docker argv used for the pass. The evidence collector produces container proof, API probe records, UI screenshots, and an evidence manifest containing hashes. Behavioural evidence is advisory to the deterministic score unless the scoring design explicitly selects it for a hidden judge; the contract says the judge sees behavioural evidence only and cannot overturn a deterministic pass.
The downstream mapping is consequently layered. Image and launch facts explain the execution substrate. Pass status and transcript facts explain whether the agent process ran. Visible and hidden score fields decide task outcome. Audit and protocol fields decide whether that outcome is valid for primary analysis. Evidence fields support review and behavioural confirmation. None of the container proof fields by themselves mean that the task passed.
12. Tests
The directly relevant test files are 5. Experiment/1. Harness/scripts/tests/test_run_cell.py, test_agent_backends.py, test_stop_and_evidence.py, test_audit_transcripts.py, test_isolation_contract.py, test_aggregate.py, and test_score_cell.py. test_run_cell.py exercises the Docker argv shape, writable configuration temporary filesystem, subscription and API route behaviour, shell availability, visible-test container command, cidfile arguments, and parent-repository fingerprint cases. It also exercises the pass-loop responses to guarded changes. test_agent_backends.py exercises the shell capability seam used by backend invocation. test_stop_and_evidence.py exercises durable cidfile proof and stop behaviour. test_audit_transcripts.py exercises path extraction and contamination classification. test_isolation_contract.py exercises the frozen contract text and registration. test_aggregate.py and test_score_cell.py exercise downstream reading of shell and scoring evidence.
The repository does not provide a reviewed mapping from every section 10 exit path to one named test. The following status is therefore deliberately conservative. Invalid billing has route and credential cases in test_run_cell.py, but the exact aggregate effect is not mapped here. Image mismatch has no named test verified in the supplied search. Workspace preparation refusal has broad runner tests, but no one-to-one mapping was verified. Container launch and shell failure have command-shape and shell capability coverage in test_run_cell.py and test_stop_and_evidence.py; the complete Docker daemon failure path was not verified. Working-tree and committed-tree changes have fingerprint coverage in test_run_cell.py; transcript contamination has coverage in test_audit_transcripts.py. Missing or malformed transcript coverage is present at the audit function level, but the full downstream driver path was not mapped. Evidence capture branches have file-level tests in test_stop_and_evidence.py and downstream evidence tests, but every server and screenshot failure branch was not mapped.
This is a statement about traceability, not a claim that an unlisted path is untested. The mapping of tests to exit paths was not made in the reviewed design evidence.
13. Dated incidents and code history
The billing route comments at run_cell.py lines 380 to 398 and 430 to 435 carry the incident behind explicit route selection. Run set 016 was launched with both authentication forms, used the metered account, and exhausted purchased credit after the earlier run sets. The source dates the corrective design by its references to run sets 014, 015, and 016. The corresponding guard is not an arbitrary refusal. It prevents a silent account change before token use.
The temporary filesystem comments at run_cell.py lines 548 to 580 carry the 2026-09-03 shell failure. The read-only credentials mount caused a root-owned configuration parent, so the unprivileged agent could not create its session directory. The comments name run sets 021, 060, 062, and 066 and the recorded review path. The unconditional temporary filesystem is the repair. The source also records that it is mounted even on the route without a credentials file so the command has one reasoning path.
The cidfile comments at run_cell.py lines 503 to 508 carry the 2026-08-29 transparency audit. Before that change, a cell artifact did not tie a pass to a container. The pass-specific cidfile and later container_evidence.json are the resulting proof chain. The flow model also records the container proof event as evidence.container_proof and describes its post-scoring construction.
The parent fingerprint comments at run_cell.py lines 678 to 697 carry the 2026-07-28 escape and the 2026-08-15 narrowing. The first incident involved an absolute-path escape that edited a frozen variant and committed to the parent repository. The earlier whole-HEAD guard then caused false stops when researchers edited plan documents during run set 009. The current guard watches only guarded subtrees, distinguishes committed-tree changes from working-tree-only changes, and fails loudly when a subtree is present but cannot be fingerprinted, at lines 707 to 729.
The isolation contract is frozen at version 1 dated 2026-09-09, as stated in its lines 1 to 10. Its history-bearing comments and rules distinguish the current maintenance agent from future roles, keep hidden scoring host-side, restrict retries to sanitized task-visible feedback, and disallow cross-cell memory. These are contract constraints rather than assumptions inferred from a diagram.
14. The weakest claim, what was not checked, and the token line
The weakest claim is that the complete campaign-wide downstream treatment of every container, transcript, audit, and evidence failure is uniform. The supplied sources verify the owned functions, records, and major branches, but they do not expose every caller’s final aggregation rule or every external Docker failure. The exact Dockerfile and shell build line ranges, the internal screenshot script, and a one-to-one test mapping for every unhappy path were not checked beyond the cited digest and searches. The token line is:
Model: openai/gpt-5.6-luna via OpenRouter; tokens: see the job ledger