Chapter 1.5 follows chapter 1.4, The eight documentation doses, and precedes chapter 1.6, The pre-registered analysis plan, within part 1, the reading path.

A cell, meaning one complete measured run of one maintenance task, moves from a disposable workspace to a recorded outcome. An arm, meaning a codebase version under comparison, supplies the version. A seed, meaning a run identifier, identifies the run. A task, meaning a maintenance change assigned to the agent, gives the agent its work. A hidden tier tests the result. A probe supplies evidence from a hidden check. A witness supports the host-side record.

1.5.1 Cell sequence of the documentation experiment

The cell scroll sequence follows staging, the act of preparing an untouched workspace, and the admission gate, the check that decides whether that workspace may begin. It then follows the sealed session, visible checks, hidden checks, and the recorded result. The harness keeps the frozen arm unchanged while it records the work in a disposable cell folder.

Figure D-H0-1. The life of one measured cell from the lease to its row in the campaign tables

%% figure D-H0-1
flowchart TB
  subgraph driver["Batch driver"]
  n1_1["step"] --> n1_2["run_batch()"]
  n1_2 --> n1_3["lane_coordinator.py"]
  n1_3 --> n1_4["LaneSpec"]
  n2_1["run_lanes()"] --> n2_2["preflight_batch.py"]
  n2_2 --> n2_3["preflight()"]
  n2_3 --> n2_4["probe_task_registrations()"]
  n3_1["<condition>__<task>__<phase>__s<seed>"] --> n3_2["step"]
  n3_2 --> n3_3["step"]
  n3_3 --> n3_4["step"]
  n4_1["stage_working_copy"] --> n4_2["hash_before"]
  n4_2 --> n4_3["snapshot_staged_tree"]
  n4_3 --> n4_4["scan_task_brief"]
  n5_1["staging_gate"] --> n5_2["step"]
  n5_2 --> n5_3["step"]
  n5_3 --> n5_4["for k in range(start_pass, max_passes + 1)"]
  n6_1["budget_exhausted"] --> n6_2["claude"]
  n6_2 --> n6_3["dgx_codex"]
  n6_3 --> n6_4["step"]
  n7_1["step"] --> n7_2["BACKEND_NAMES"]
  n7_2 --> n7_3["step"]
  n7_3 --> n7_4["cid_path"]
  n8_1["dgx_fabric.py"] --> n8_2["FabricClient"]
  n8_2 --> n8_3["step"]
  n8_3 --> n8_4["_FabricLeaseBackend"]
  end
  subgraph records["Cell records"]
  n9_1["capture_handoff"] --> n9_2["handoff_lib.py"]
  n9_2 --> n9_3["validate()"]
  n9_3 --> n9_4["1. Harness/config/agent_handoff.schema.json"]
  n10_1["build_token_ledger"] --> n10_2["ledger_lib.py"]
  n10_2 --> n10_3["build_ledger()"]
  n10_3 --> n10_4["repo_escape_check"]
  n11_1["parent_repo_fingerprint()"] --> n11_2["visible_tests"]
  n11_2 --> n11_3["step"]
  n11_3 --> n11_4["make_patch"]
  n12_1["hidden_checks"] --> n12_2["step"]
  n12_2 --> n12_3["run_hidden_scoring()"]
  n12_3 --> n12_4["compose_retry_notice()"]
  n13_1["unmet_criterion()"] --> n13_2["retry_feedback.json"]
  n13_2 --> n13_3["remove_workspace_git"]
  n13_3 --> n13_4["hash_after"]
  n14_1["_finalize_manifests()"] --> n14_2["cell_manifest.json"]
  n14_2 --> n14_3["timings.jsonl"]
  n14_3 --> n14_4["hidden_tier_result.json"]
  n15_1["transcripts/"] --> n15_2["patches/"]
  n15_2 --> n15_3["scoring_outputs/"]
  n15_3 --> n15_4["step"]
  n16_1["score_cell()"] --> n16_2["collect_invalid_reasons()"]
  n16_2 --> n16_3["6. Metrics/cell_factor_matrix.csv"]
  n16_3 --> n16_4["run_id"]
  n17["cell_id"]
  n1_4 --> n2_1
  n2_4 --> n3_1
  n3_4 --> n4_1
  n4_4 --> n5_1
  n5_4 --> n6_1
  n6_4 --> n7_1
  n7_4 --> n8_1
  n8_4 --> n9_1
  n9_4 --> n10_1
  n10_4 --> n11_1
  n11_4 --> n12_1
  n12_4 --> n13_1
  n13_4 --> n14_1
  n14_4 --> n15_1
  n15_4 --> n16_1
  end
  n16_4 --> n17
  n1_1 -- "run_batch.py" --> n1_1_reference["reference"]
  n3_2 -- "run_batch.py" --> n3_2_reference["reference"]
  n3_3 -- "_cell_id()" --> n3_3_reference["reference"]
  n3_4 -- "run_cell.py" --> n3_4_reference["reference"]
  n5_2 -- "run_staging_gate()" --> n5_2_reference["reference"]
  n5_3 -- "run_cell.py" --> n5_3_reference["reference"]
  n6_4 -- "dgx_claude" --> n6_4_reference["reference"]
  n7_1 -- "agent_backends.py" --> n7_1_reference["reference"]
  n7_3 -- "run_cell.py" --> n7_3_reference["reference"]
  n8_3 -- "agent_backends.py" --> n8_3_reference["reference"]
  n11_3 -- "visible_tests_host_confirm" --> n11_3_reference["reference"]
  n12_2 -- "score_cell.py" --> n12_2_reference["reference"]
  n15_4 -- "score_cell.py" --> n15_4_reference["reference"]

source: operations/site-ia/A3-diagram-specification.md source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”

Cell scroll sequence from staging to recorded closeout.

1.5.2 Staging workspace of the documentation experiment

The staging workspace copies an untouched frozen arm into an isolated workspace and gives that copy a repository boundary for cells in the workspace. The workspace receives the seeded task and the agent works only there. The parent research repository remains outside the workspace.

Figure D-R4-6. The disposable staging workspace, with the frozen arm copied into an isolated cell workspace before the agent session.

%% figure D-R4-6
flowchart LR
  n1["The disposable staging workspace<br/>the frozen arm copied into an isolated cell workspace before the agent session"]

source: operations/site-ia/A3-diagram-specification.md source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”

Disposable staging workspace before admission.

1.5.3 Admission gate of the documentation experiment

The admission gate is the starting-tree check that decides whether a cell may begin. The hidden tier runs against the staged tree. A cell proceeds when the starting tree fails the required hidden checks and the task has a registered scorer. A tree that already passes, or a task without a hidden scorer, receives a substrate-absent disposition.

1.5.4 Sealed container of the documentation experiment

A container, a sealed execution environment holding the workspace and the agent process, runs the session. A lease, a reserved serving slot assigned to one cell and tied to a serving controller, binds the session to its serving resource. The sealed boundary keeps hidden checks and their output outside the agent workspace.

Figure D-R4-7. The sealed container boundary, with the agent workspace inside the container and host-side checks outside it.

%% figure D-R4-7
flowchart LR
  n1["The sealed container boundary<br/>the agent workspace inside the container and host-side checks outside it"]

source: operations/site-ia/A3-diagram-specification.md source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”

Sealed container boundary for the cell session.

1.5.5 Request protocols and routes

The two request protocols are the OpenAI Responses protocol and the Anthropic Messages protocol used to request model work. A paid route is the hosted model path used for comparison. A local route is a path from the harness to local inference. The two local routes are dgx_codex and dgx_claude. The data center GPU (DGX) Spark is the local inference machine. Fabric is the model-serving controller that assigns leases. The agent-coder is the alias of the local coding model.

The local routes support connectivity shakedowns and are invalid for primary analysis. The paid route supplies the comparison path.

1.5.6 Pass loop of the documentation experiment

A pass is one agent attempt followed by visible and hidden checks. A failed pass returns a fixed failure note to the agent. Tokens to success are the cumulative tokens spent before the cell reaches a passing disposition. The stopping rule is the condition that ends the loop at success, the pass ceiling, or the token ceiling.

The pass limit is five. source: 1. Harness/config/run_defaults.json line 3 (key max_passes)

The token budget to success is 2,500,000 tokens, cumulative across passes. source: 1. Harness/config/run_defaults.json line 4 (key token_budget_per_cell) source: 5. Experiment/0. Plan/ section (d) line 112

The measured local route has a budget of 30,000,000 tokens. Its basis is a measured run that used 8,240,883 ledger tokens for one 95-turn pass. source: run_defaults.json line 6 (key token_budget_per_cell_dgx_claude); basis at line 7, recording run set 035 seed 2

The flaky re-run count is three when the agent-visible result disagrees with the harness-visible result. source: run_defaults.json line 9 (key flaky_rerun_n)

1.5.7 Token categories of the documentation experiment

The ledger is the record of token usage by category. A tranche is one recorded allocation of work or scoring time. A replay hash is the digest used to check that a recorded input and output can be replayed. The ledger has ten run categories and the scoring-only category scoring_external.

Figure D-R4-2. The token categories in the cell ledger, separating agent work, checks, scoring, and the scoring-only external category.

%% figure D-R4-2
flowchart LR
  n1["The token categories in the cell ledger<br/>separating agent work<br/>checks<br/>scoring"]
  n2["the scoring-only external category"]
  n1 --> n2

source: operations/site-ia/A3-diagram-specification.md source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”

Token categories recorded by the cell ledger.

source: 1. Harness/scripts/ledger_lib.py lines 55 to 59

1.5.8 Closeout record of the documentation experiment

Closeout, the fixed record chain that files the cell after its session and scoring, preserves the final evidence. Disposition, the page-level outcome assigned to that record, describes how the cell ended. The four score states are pass, fail, not_run, and invalid. The scoring contract defines seven invalidity conditions.

source: 1. Harness/config/scoring_contract.json keys score_states and invalid_if

Figure D-R4-3. The outcome branches from a completed cell record to success, ceiling failure, contamination, substrate absence, or invalid record.

%% figure D-R4-3
flowchart LR
  n1["The outcome branches from a completed cell record to success<br/>ceiling failure<br/>contamination<br/>substrate absence"]
  n2["or invalid record"]
  n1 --> n2

source: operations/site-ia/A3-diagram-specification.md source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”

Outcome branches for the cell disposition.

Page dispositionContract stateDistinguishing record or fieldScorer path or decision source
Successpassvisible_pass and hidden_pass are true, with no invalid reasonsscore_cell() computes correctness and sets overall to pass
Failure at the pass ceilingfailmetrics.json budget_exhausted is true and passes_run reaches the configured pass limitThe runner records the ceiling before scoring; score_cell() sets overall to fail when no invalid reason applies
Failure at the token ceilingfailmetrics.json budget_exhausted is true and the token ledger reaches the configured token budget before the pass limitThe runner records the ceiling before scoring; score_cell() sets overall to fail when no invalid reason applies
Contaminatedinvalidinvalid_reasons contains an instrument or run-halt reason, including shell_unavailable or run_halted:<reason>collect_invalid_reasons() assembles the reasons and score_cell() sets overall to invalid
Substrate absentnot_runNo scorer record is created because the required substrate or start tree is absentThe runner or batch record records the not-run reason outside score_cell.py
Invalid recordinvalidinvalid_reasons contains a missing required record such as missing_visible_test_output, missing_hidden_scoring_output, missing_token_usage or missing_reproducibility_manifestcollect_invalid_reasons() assembles the reasons and score_cell() sets overall to invalid

1.5.9 Budget controls of the documentation experiment

The runner applies the pass ceiling and token ceiling before closeout. The measured local route has its separate token budget because its provider reports no cache fields. The closeout record preserves the stopping decision, token ledger, timing record, visible result, hidden result, and reproducibility material.

Figure D-R4-5. The extra stopping rule, showing the token and pass ceilings as independent conditions that end a cell.

%% figure D-R4-5
flowchart LR
  n1["The extra stopping rule<br/>the token and pass ceilings as independent conditions that end a cell"]

source: operations/site-ia/A3-diagram-specification.md source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”

Independent pass and token stopping conditions.

1.5.10 Worked cell record of the documentation experiment

The worked cell is H-NON__H-M020__full__s8 in run set 195-20260912-auto-dgx-agent-coder-0425. Its closeout record reports complete, and its scoring summary reports an overall pass that is valid for primary analysis. The cell was admitted, completed after one pass, passed both tiers, and spent 2,145,302 tokens in total. The ledger divides that total into 1,850,000 tokens for the agent pass, 200,000 tokens for visible tests, and 95,302 tokens for hidden probes and full-suite regression.

source: cell folder H-NON__H-M020__full__s8 in run set 195-20260912-auto-dgx-agent-coder-0425, from token_ledger.jsonl and timings.jsonl

Figure D-R4-4. The worked cell timeline, generated from the cell timings record and token ledger, showing staging, admission, agent work, visible checks, hidden checks, and final disposition.

%% figure D-R4-4
flowchart LR
  n1["The worked cell timeline<br/>generated from the cell timings record and token ledger<br/>staging<br/>admission"]
  n2["agent work<br/>visible checks<br/>hidden checks<br/>final disposition"]
  n1 --> n2

source: operations/site-ia/A3-diagram-specification.md source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”

Worked cell timeline from the timing record and token ledger.

The worked cell is a shakedown record set apart from confirmatory analysis. Its ledger and timing files show how the harness records a complete measured cell.

Previous: The eight documentation doses | Next: The pre-registered analysis plan
source: `operations/site-ia/A6-page-briefs.md` section "### R4. Measurement of one cell"