Chapter 1.5 follows chapter 1.4, The eight documentation doses, and precedes chapter 1.6, The pre-registered analysis plan, within part 1, the reading path.
A cell, meaning one complete measured run of one maintenance task, moves from a disposable workspace to a recorded outcome. An arm, meaning a codebase version under comparison, supplies the version. A seed, meaning a run identifier, identifies the run. A task, meaning a maintenance change assigned to the agent, gives the agent its work. A hidden tier tests the result. A probe supplies evidence from a hidden check. A witness supports the host-side record.
1.5.1 Cell sequence of the documentation experiment
The cell scroll sequence follows staging, the act of preparing an untouched workspace, and the admission gate, the check that decides whether that workspace may begin. It then follows the sealed session, visible checks, hidden checks, and the recorded result. The harness keeps the frozen arm unchanged while it records the work in a disposable cell folder.
Figure D-H0-1. The life of one measured cell from the lease to its row in the campaign tables
%% figure D-H0-1 flowchart TB subgraph driver["Batch driver"] n1_1["step"] --> n1_2["run_batch()"] n1_2 --> n1_3["lane_coordinator.py"] n1_3 --> n1_4["LaneSpec"] n2_1["run_lanes()"] --> n2_2["preflight_batch.py"] n2_2 --> n2_3["preflight()"] n2_3 --> n2_4["probe_task_registrations()"] n3_1["<condition>__<task>__<phase>__s<seed>"] --> n3_2["step"] n3_2 --> n3_3["step"] n3_3 --> n3_4["step"] n4_1["stage_working_copy"] --> n4_2["hash_before"] n4_2 --> n4_3["snapshot_staged_tree"] n4_3 --> n4_4["scan_task_brief"] n5_1["staging_gate"] --> n5_2["step"] n5_2 --> n5_3["step"] n5_3 --> n5_4["for k in range(start_pass, max_passes + 1)"] n6_1["budget_exhausted"] --> n6_2["claude"] n6_2 --> n6_3["dgx_codex"] n6_3 --> n6_4["step"] n7_1["step"] --> n7_2["BACKEND_NAMES"] n7_2 --> n7_3["step"] n7_3 --> n7_4["cid_path"] n8_1["dgx_fabric.py"] --> n8_2["FabricClient"] n8_2 --> n8_3["step"] n8_3 --> n8_4["_FabricLeaseBackend"] end subgraph records["Cell records"] n9_1["capture_handoff"] --> n9_2["handoff_lib.py"] n9_2 --> n9_3["validate()"] n9_3 --> n9_4["1. Harness/config/agent_handoff.schema.json"] n10_1["build_token_ledger"] --> n10_2["ledger_lib.py"] n10_2 --> n10_3["build_ledger()"] n10_3 --> n10_4["repo_escape_check"] n11_1["parent_repo_fingerprint()"] --> n11_2["visible_tests"] n11_2 --> n11_3["step"] n11_3 --> n11_4["make_patch"] n12_1["hidden_checks"] --> n12_2["step"] n12_2 --> n12_3["run_hidden_scoring()"] n12_3 --> n12_4["compose_retry_notice()"] n13_1["unmet_criterion()"] --> n13_2["retry_feedback.json"] n13_2 --> n13_3["remove_workspace_git"] n13_3 --> n13_4["hash_after"] n14_1["_finalize_manifests()"] --> n14_2["cell_manifest.json"] n14_2 --> n14_3["timings.jsonl"] n14_3 --> n14_4["hidden_tier_result.json"] n15_1["transcripts/"] --> n15_2["patches/"] n15_2 --> n15_3["scoring_outputs/"] n15_3 --> n15_4["step"] n16_1["score_cell()"] --> n16_2["collect_invalid_reasons()"] n16_2 --> n16_3["6. Metrics/cell_factor_matrix.csv"] n16_3 --> n16_4["run_id"] n17["cell_id"] n1_4 --> n2_1 n2_4 --> n3_1 n3_4 --> n4_1 n4_4 --> n5_1 n5_4 --> n6_1 n6_4 --> n7_1 n7_4 --> n8_1 n8_4 --> n9_1 n9_4 --> n10_1 n10_4 --> n11_1 n11_4 --> n12_1 n12_4 --> n13_1 n13_4 --> n14_1 n14_4 --> n15_1 n15_4 --> n16_1 end n16_4 --> n17 n1_1 -- "run_batch.py" --> n1_1_reference["reference"] n3_2 -- "run_batch.py" --> n3_2_reference["reference"] n3_3 -- "_cell_id()" --> n3_3_reference["reference"] n3_4 -- "run_cell.py" --> n3_4_reference["reference"] n5_2 -- "run_staging_gate()" --> n5_2_reference["reference"] n5_3 -- "run_cell.py" --> n5_3_reference["reference"] n6_4 -- "dgx_claude" --> n6_4_reference["reference"] n7_1 -- "agent_backends.py" --> n7_1_reference["reference"] n7_3 -- "run_cell.py" --> n7_3_reference["reference"] n8_3 -- "agent_backends.py" --> n8_3_reference["reference"] n11_3 -- "visible_tests_host_confirm" --> n11_3_reference["reference"] n12_2 -- "score_cell.py" --> n12_2_reference["reference"] n15_4 -- "score_cell.py" --> n15_4_reference["reference"]
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”
Cell scroll sequence from staging to recorded closeout.
1.5.2 Staging workspace of the documentation experiment
The staging workspace copies an untouched frozen arm into an isolated workspace and gives that copy a repository boundary for cells in the workspace. The workspace receives the seeded task and the agent works only there. The parent research repository remains outside the workspace.
Figure D-R4-6. The disposable staging workspace, with the frozen arm copied into an isolated cell workspace before the agent session.
%% figure D-R4-6 flowchart LR n1["The disposable staging workspace<br/>the frozen arm copied into an isolated cell workspace before the agent session"]
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”
Disposable staging workspace before admission.
1.5.3 Admission gate of the documentation experiment
The admission gate is the starting-tree check that decides whether a cell may begin. The hidden tier runs against the staged tree. A cell proceeds when the starting tree fails the required hidden checks and the task has a registered scorer. A tree that already passes, or a task without a hidden scorer, receives a substrate-absent disposition.
1.5.4 Sealed container of the documentation experiment
A container, a sealed execution environment holding the workspace and the agent process, runs the session. A lease, a reserved serving slot assigned to one cell and tied to a serving controller, binds the session to its serving resource. The sealed boundary keeps hidden checks and their output outside the agent workspace.
Figure D-R4-7. The sealed container boundary, with the agent workspace inside the container and host-side checks outside it.
%% figure D-R4-7 flowchart LR n1["The sealed container boundary<br/>the agent workspace inside the container and host-side checks outside it"]
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”
Sealed container boundary for the cell session.
1.5.5 Request protocols and routes
The two request protocols are the OpenAI Responses protocol and the Anthropic Messages protocol used to request model work. A paid route is the hosted model path used for comparison. A local route is a path from the harness to local inference. The two local routes are dgx_codex and dgx_claude. The data center GPU (DGX) Spark is the local inference machine. Fabric is the model-serving controller that assigns leases. The agent-coder is the alias of the local coding model.
The local routes support connectivity shakedowns and are invalid for primary analysis. The paid route supplies the comparison path.
1.5.6 Pass loop of the documentation experiment
A pass is one agent attempt followed by visible and hidden checks. A failed pass returns a fixed failure note to the agent. Tokens to success are the cumulative tokens spent before the cell reaches a passing disposition. The stopping rule is the condition that ends the loop at success, the pass ceiling, or the token ceiling.
The pass limit is five.
source: 1. Harness/config/run_defaults.json line 3 (key max_passes)
The token budget to success is 2,500,000 tokens, cumulative across passes.
source: 1. Harness/config/run_defaults.json line 4 (key token_budget_per_cell)
source: 5. Experiment/0. Plan/ section (d) line 112
The measured local route has a budget of 30,000,000 tokens. Its basis is a measured run that used 8,240,883 ledger tokens for one 95-turn pass.
source: run_defaults.json line 6 (key token_budget_per_cell_dgx_claude); basis at line 7, recording run set 035 seed 2
The flaky re-run count is three when the agent-visible result disagrees with the harness-visible result.
source: run_defaults.json line 9 (key flaky_rerun_n)
1.5.7 Token categories of the documentation experiment
The ledger is the record of token usage by category. A tranche is one recorded allocation of work or scoring time. A replay hash is the digest used to check that a recorded input and output can be replayed. The ledger has ten run categories and the scoring-only category scoring_external.
Figure D-R4-2. The token categories in the cell ledger, separating agent work, checks, scoring, and the scoring-only external category.
%% figure D-R4-2 flowchart LR n1["The token categories in the cell ledger<br/>separating agent work<br/>checks<br/>scoring"] n2["the scoring-only external category"] n1 --> n2
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”
Token categories recorded by the cell ledger.
source: 1. Harness/scripts/ledger_lib.py lines 55 to 59
1.5.8 Closeout record of the documentation experiment
Closeout, the fixed record chain that files the cell after its session and scoring, preserves the final evidence. Disposition, the page-level outcome assigned to that record, describes how the cell ended. The four score states are pass, fail, not_run, and invalid. The scoring contract defines seven invalidity conditions.
source: 1. Harness/config/scoring_contract.json keys score_states and invalid_if
Figure D-R4-3. The outcome branches from a completed cell record to success, ceiling failure, contamination, substrate absence, or invalid record.
%% figure D-R4-3 flowchart LR n1["The outcome branches from a completed cell record to success<br/>ceiling failure<br/>contamination<br/>substrate absence"] n2["or invalid record"] n1 --> n2
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”
Outcome branches for the cell disposition.
| Page disposition | Contract state | Distinguishing record or field | Scorer path or decision source |
|---|---|---|---|
| Success | pass | visible_pass and hidden_pass are true, with no invalid reasons | score_cell() computes correctness and sets overall to pass |
| Failure at the pass ceiling | fail | metrics.json budget_exhausted is true and passes_run reaches the configured pass limit | The runner records the ceiling before scoring; score_cell() sets overall to fail when no invalid reason applies |
| Failure at the token ceiling | fail | metrics.json budget_exhausted is true and the token ledger reaches the configured token budget before the pass limit | The runner records the ceiling before scoring; score_cell() sets overall to fail when no invalid reason applies |
| Contaminated | invalid | invalid_reasons contains an instrument or run-halt reason, including shell_unavailable or run_halted:<reason> | collect_invalid_reasons() assembles the reasons and score_cell() sets overall to invalid |
| Substrate absent | not_run | No scorer record is created because the required substrate or start tree is absent | The runner or batch record records the not-run reason outside score_cell.py |
| Invalid record | invalid | invalid_reasons contains a missing required record such as missing_visible_test_output, missing_hidden_scoring_output, missing_token_usage or missing_reproducibility_manifest | collect_invalid_reasons() assembles the reasons and score_cell() sets overall to invalid |
1.5.9 Budget controls of the documentation experiment
The runner applies the pass ceiling and token ceiling before closeout. The measured local route has its separate token budget because its provider reports no cache fields. The closeout record preserves the stopping decision, token ledger, timing record, visible result, hidden result, and reproducibility material.
Figure D-R4-5. The extra stopping rule, showing the token and pass ceilings as independent conditions that end a cell.
%% figure D-R4-5 flowchart LR n1["The extra stopping rule<br/>the token and pass ceilings as independent conditions that end a cell"]
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”
Independent pass and token stopping conditions.
1.5.10 Worked cell record of the documentation experiment
The worked cell is H-NON__H-M020__full__s8 in run set 195-20260912-auto-dgx-agent-coder-0425. Its closeout record reports complete, and its scoring summary reports an overall pass that is valid for primary analysis. The cell was admitted, completed after one pass, passed both tiers, and spent 2,145,302 tokens in total. The ledger divides that total into 1,850,000 tokens for the agent pass, 200,000 tokens for visible tests, and 95,302 tokens for hidden probes and full-suite regression.
source: cell folder H-NON__H-M020__full__s8 in run set 195-20260912-auto-dgx-agent-coder-0425, from token_ledger.jsonl and timings.jsonl
Figure D-R4-4. The worked cell timeline, generated from the cell timings record and token ledger, showing staging, admission, agent work, visible checks, hidden checks, and final disposition.
%% figure D-R4-4 flowchart LR n1["The worked cell timeline<br/>generated from the cell timings record and token ledger<br/>staging<br/>admission"] n2["agent work<br/>visible checks<br/>hidden checks<br/>final disposition"] n1 --> n2
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section ”# R4. Measurement of one cell”
Worked cell timeline from the timing record and token ledger.
The worked cell is a shakedown record set apart from confirmatory analysis. Its ledger and timing files show how the harness records a complete measured cell.