Sample chapter of the detailed design, written 2026-09-12 by Claude Fable 5.1 to the standard set in 5. Experiment/11. Detailed Design/specifications/C0-detailed-design-specification.md. Every line number below was read from 5. Experiment/1. Harness/scripts/run_cell.py at commit b69ae5977 (the file last changed in commit beb00f575 of 2026-09-09), and from the neighbouring files named beside each reference at the same commit. The generated digest 5. Experiment/11. Detailed Design/records/C-code-digest.md was the map; it reports every function of this file as spanning lines 1 to 3031, so no line number here comes from the digest. The graph under operations/detailed-design/graphify-out/ places the file’s functions in six communities (11, 45, 48, 52, 65 and 100 in GRAPH_REPORT.md), and its path tool confirms that the runner and the scorer meet only through the shared timing module (graphify path "run()" "score_cell()" answers with a two-hop path through StageTimer).

Owning files. 5. Experiment/1. Harness/scripts/run_cell.py (3,265 lines, 166,731 bytes). Files this chapter reads into but does not own: agent_backends.py (the lease-holding back ends, chapter 05), dgx_fabric.py (the controller client, chapter 03), score_cell.py (the hidden scoring it borrows during a pass, chapter 08), timing_lib.py and ledger_lib.py (the two record writers, chapter 09), handoff_lib.py, lint_task_briefs.py, publish_lock.py, and run_batch.py (the caller, chapter 02).

1. What a cell is and what the runner does

A cell is one measured attempt: one coding agent, one maintenance task, one product build, one repetition seed, run at most five times within one cumulative token budget. The cell runner is the program that performs one cell from beginning to end. It takes a frozen copy of the product build, prepares it so that the task is genuinely undone, lets the agent try, tests what the agent did, and writes the records that every later program reads. Nothing downstream ever talks to an agent; the runner is the only program in the harness that spends tokens on a measured agent, and everything the experiment later says about a cell is read from the files the runner leaves behind.

The runner is invoked once per cell by the batch driver, run_batch.py, through a subprocess (run_batch.py line 113, default_invoke; the arguments are assembled at lines 179 to 216, _run_cell_args). It may be invoked a second time for the same cell with the flag --resume-hidden-fail, which continues the same cell rather than starting a new one (section 4.3). It writes into one cell directory, 5. Experiment/5. Run Sets/<run set>/cells/<cell id>/, and appends to three files at the run-set level. Its result is a Python dictionary printed as one JSON line on standard output (run_cell.py lines 3255 to 3262, main) and an exit code that is 0 for every outcome except a refusal, which returns 2.

2. The reader’s map of the file

The file has five regions. Lines 1 to 21 are the module docstring, which still describes the July design (the Claude command line and a static context-window map) and does not mention the Fabric back ends added in August. Lines 23 to 59 are the imports, which name every sibling module the runner depends on. Lines 61 to 270 are constants, and most of the experiment’s rules about a cell live here as literal text: the single prompt template and its retired predecessors (lines 104 to 107), the prompt variants (lines 121 to 131), the context-window map (line 140), the sets of directory and file names that are excluded from every hash, snapshot and patch (lines 146 to 169), the three guarded subtrees the escape guard watches (line 176), and the four verbatim retry notices the agent may be shown (lines 199 to 268). Lines 271 to 1673 are helper functions, each of which is described in the section that uses it. Lines 1674 to 3029 are the main function, run, which is one function of 1,356 lines with two nested functions, _refuse (line 1859) and _satisfied_by_start_tree (line 1870). Lines 3032 to 3265 are the finalisation helpers, the signal handler and the entry point.

The main function is written as numbered steps in comments: step 1 resolve and validate (line 1767), step 2 stage the working copy (line 1959), step 3 inject the realisation (line 2079), step 3b the staging gate (line 2162), step 4 snapshot the manifests (line 2203), step 5 the pass loop (line 2361), and step 6 finalise (line 2921). Section 4 of this chapter follows the same order.

3. The inputs

3.1 Command-line arguments

The parser is _parse_args at lines 1508 to 1557. Required: --run-set (the run-set folder, absolute or relative to the base), --project, --condition (the build, for example H-LAP-L5P or H-STR), --task (for example H-M024), and --seed (an integer). Optional: --phase (defaults to full), --model, --prompt-variant (none or artifact_hint, amendment A8), --retry-feedback (on or off, amendment A9), --agent-backend (claude, dgx_claude or dgx_codex), --max-passes, --token-budget, --dry-run, --resume-hidden-fail, --base-dir, --fabric-endpoint (dgx or igx_thor), and --start-tree (a previous cell’s finished workspace, for a step of a chained sequence; filled only by the batch driver).

3.2 Environment variables

VariableRead atWhat it decides
FIVEB_CELL_CONTAINERline 1705The container image; unset means the agent and the visible tests run on the host. Required for any dgx_* back end (line 1714).
FIVEB_BILLING, FIVEB_CELL_CREDENTIALS, ANTHROPIC_API_KEYbilling_route lines 375 to 425, assert_billing_ready lines 427 to 449, and the refusal at lines 1741 to 1749Which account pays for an Anthropic-route cell, and the refusal of any Anthropic credential on the dgx_claude route.
FIVEB_FABRIC_ENDPOINTline 1722Which physical controller a Fabric cell admits against, when --fabric-endpoint is absent.
DGX_SPARK_FABRIC_MODEL, DGX_SPARK_FABRIC_LANE, DGX_SPARK_FABRIC_PRIORITY, DGX_SPARK_FABRIC_LEASE_SECONDS, DGX_SPARK_FABRIC_REQUEST_PRESET, DGX_SPARK_FABRIC_POLL_SECONDSlines 1682, 1725, and the fabric block of the manifest at lines 2322 to 2340, line 2382The Fabric admission request: alias, lane, priority, lease length, pinned request preset, and the polling interval.
FIVEB_LANE_OWNER, FIVEB_LANE_ENDPOINT, FIVEB_LANE_RUN_IDline 1797 and lane_container_labels lines 344 to 365The lane identity a coordinator-launched cell inherits; it becomes Docker labels on every container the cell starts.
FIVEB_STAGE_NODE_MODULESline 2010Set to 0 to exclude the JavaScript dependency tree from the staged copy.
FIVEB_SKIP_STAGING_GATEline 2167Set to 1 to skip the staging gate; the manifest records the skip.
FIVEB_CHAINED_GATE_SATISFIEDchained_gate_ruling_enabled lines 946 to 958Set to 0 to restore the pre-ruling refusal of a chained step whose start tree already passes.
FIVEB_LEASE_RETRY_SECONDSagent_backends.py line 113The wait between lease attempts while the controller cannot grant one; default 60 seconds.

3.3 Files read

The runner reads the run defaults config/run_defaults.json and the model matrix config/model_matrix.json (lines 1678 to 1679), the container digest file container/image_digest.txt (line 1751, parsed by _read_container_digest_file lines 1026 to 1038), the single prompt template prompts/maintenance.single_agent.v006.md (line 1788) and, under the hint variant, prompts/artifact_hint.v001.md (compose_artifact_hint lines 869 to 922). From the task folder it reads task_manifest.json (line 1826), the agent-facing pair task.md and visible_tests.md or, in the current suite, TICKET.md and HOW_TO_VERIFY.md (compose_task_package lines 1560 to 1581), the realisation diffs named by the manifest (_resolve_realization_diff lines 1246 to 1257) or the feature-removal diffs reference_feature.<condition>.diff and reference_feature.diff (lines 2143 to 2146), and the frozen retry-feedback map retry_feedback.json (load_retry_feedback lines 271 to 288). From the build it reads the whole variant tree (line 2055) and the artefact inventory lap_artifact_manifest.json (line 2209 and ledger_lib.load_lap_manifest line 212). On a resume it reads its own earlier metrics.json, cell_manifest.json, token_ledger.jsonl or provider_usage.jsonl, and the transcript names (lines 1845 to 1857, 2075 to 2079, 2406 to 2435).

3.4 Network

An Anthropic-route cell reaches the Anthropic service from inside the container through the agent’s own command-line tool; the runner makes no network call itself. A Fabric-route cell reaches the controller of the DGX Spark or the IGX Thor through dgx_fabric.FabricClient (chapter 03): admission, polling, claim, heartbeat, release and cancellation (dgx_fabric.py lines 381 to 450 and 525 to 563). The runner also runs git and docker as subprocesses on the host (parent_repo_fingerprint lines 675 to 743, runtime_image_identity lines 1041 to 1064).

4. The happy path in order

4.1 Resolution and validation, lines 1675 to 1790

select_backend (agent_backends.py line 373) validates the back-end name. The model is the argument, then the Fabric environment alias, then the back end’s default (lines 1681 to 1685). The pass ceiling is the argument or the run default of five (line 1686). The token budget is the argument, or 30,000,000 for dgx_claude because the local Anthropic-shaped endpoint reports no cache fields and bills the whole growing context as plain input on every turn, or 2,500,000 otherwise (lines 1694 to 1700; the basis is quoted in run_defaults.json). A Fabric back end must have a container image, a resolvable endpoint, and a maintenance lane in the allowed set (lines 1713 to 1733), and on the dgx_claude route it must have no Anthropic credential in the environment, because such a credential could be reached from inside the container and silently bill an Anthropic account, which is what emptied the credit during run set 016 (lines 1734 to 1749). On the Anthropic container route the billing route is checked before any token is spent (line 1750 to 1752). The container’s identity is recorded from the build-time digest file and from the running Docker daemon (lines 1753 to 1766). The run set, task folder and variant root are resolved (lines 1769 to 1783).

4.2 Identity, the timer and the result skeleton, lines 1791 to 1831

The cell identifier is <condition>__<task>__<phase>__s<seed> (line 1791); the model is deliberately not part of it, which is why two cells differing only by model must sit in different run sets (section 4.4). The cell directory and its three subfolders transcripts, patches and scoring_outputs are created (lines 1799 to 1801). A StageTimer (timing_lib.py line 88) is constructed at line 1806; from here on every named activity appends one line to timings.jsonl and keeps a live marker timing_current.json naming the activity in progress. The result dictionary (lines 1808 to 1824) starts with status unset, and the two exploratory fields that mark a dgx_codex cell as invalid by intent (validation_purpose equal to dgx_maintenance_connectivity).

4.3 The resume decision and sticky invalidity, lines 1833 to 1857

A resume is recognised only when the flag is set and the cell already has a workspace and at least one transcript (lines 1833 to 1836). A resumed cell keeps its workspace, its staged snapshot, its ledger and its pass numbering; it is the same measured cell, and its budget is cumulative. If the earlier attempt was already invalid, that verdict is carried forward (lines 1845 to 1857), because before 2026-08-15 the result was rebuilt from scratch on every invocation and run set 009’s H-LAP-L5P__H-M018__full__s2, which had tripped the escape guard at pass 2, ended up recorded as completed and valid.

4.4 The folder-reuse guard, lines 1953 to 1958

check_cell_manifest_conflict (lines 1584 to 1671) reads any earlier cell_manifest.json in the folder and raises SystemExit with one of five reasons (cell_model_conflict, cell_fabric_endpoint_conflict, cell_fabric_endpoint_unverified, cell_prompt_variant_conflict, cell_retry_feedback_conflict) when the invocation would splice a differently configured attempt into an existing cell. It writes nothing on purpose: a refusal that rewrote the earlier cell’s status would do the damage it exists to prevent. The batch driver refuses a mixed-model batch before it starts for the same reason (run_batch.py lines 933 to 939).

4.5 Staging the working copy, lines 1961 to 2066

Unless resuming, the workspace is removed and copied afresh from the frozen build, or from the --start-tree of a chained step (lines 2031 to 2034; a missing or empty start tree is refused at line 2034). The copy keeps symbolic links and skips only .git (and node_modules when the environment says so); the comment at lines 1968 to 2013 records the measurements behind staging the dependency tree (0.91 seconds on local disk against 98 seconds on the old volume). A chained step’s start tree carries its predecessor’s HANDOFF.json in, and the frozen build’s root copy of that name is dropped (lines 2040 to 2052, amendment A12); if the note is present its hash is recorded (lines 2057 to 2061). _sha256_tree (lines 1120 to 1133) hashes every source file and writes hashes_before.json (line 2064).

4.6 Injection, lines 2079 to 2160

The task brief is scanned for study disclosure before anything is spent (lint_task_briefs.scan_task_dir, line 2089); one hit refuses the cell. A seeded-defect task applies its per-build realisation diff with _apply_diff (lines 1260 to 1279, which dry-runs the patch first so a wrong-tree diff can never half-apply), or records native_defect when the control build never implemented the behaviour (lines 2104 to 2131). A feature task removes its feature by applying the per-build diff produced by materialize_feature_bases.py or the generic reference diff (lines 2134 to 2158). Each injection is timed under inject_seeded_defect or inject_feature_removal and recorded in the manifest.

4.7 The staging gate, lines 2162 to 2186

run_staging_gate (lines 961 to 1019) runs the task’s hidden checks against the staged tree through the scorer’s own run_hidden_scoring (score_cell.py line 553). The checks must fail: a tree that already satisfies the task is a staging defect. The raw output is written to scoring_outputs/staging_gate.txt (line 2175). Three verdicts are possible. refuse ends the cell with staging_gate_failed:<detail> (line 2179), where the detail is hidden_scorer_missing or start_tree_passes_hidden. satisfied, which only a step of a chained sequence can receive and only while the ruling switch is on, hands over to _satisfied_by_start_tree (line 2186): the step is recorded as reached at zero cost, with a scoring summary (lines 1884 to 1931), a full metrics record with every count at zero (reached_step_metrics lines 3069 to 3119, written at line 1932), a REACHED line in the execution log (line 1943), and the run manifests. ok admits the agent.

4.8 The private Git boundary and the staged snapshot, lines 2189 to 2201

_initialize_workspace_git (lines 1368 to 1386) creates a repository inside the workspace so that an agent’s git add or git commit cannot reach the research repository that contains the run set. The post-injection tree is then copied to staged_snapshot/ (line 2197), which is the reference every per-pass patch is taken against.

4.9 The manifests and the prompt, lines 2203 to 2360

The artefact inventory is frozen into lap_path_manifest_snapshot.json (line 2214) in the shape the aggregator reads back (aggregate_metrics.load_snapshot, line 152). The prompt template is copied into the run set’s prompts_snapshot/ (line 2222). Under the artifact_hint variant, compose_artifact_hint (lines 869 to 922) lists the documentation files the staged copy actually carries, the composed paragraph is checked with the same disclosure lint as a brief (lines 2235 to 2242), appended to the template, and snapshotted per build (lines 2243 to 2246). The task package is composed (line 2251) and the registered visible test command found (_visible_test_command, lines 1240 to 1243). The cell manifest (lines 2254 to 2350) is written at line 2352; its fields are listed in section 7. A dry run stops here with status dry_run (lines 2356 to 2360).

4.10 The pass loop, lines 2361 to 2915

Before the loop: the token ledger and provider-usage files are removed unless resuming (lines 2363 to 2368); the Fabric back end object is built (lines 2372 to 2393, DgxCodexBackend or DgxClaudeBackend, agent_backends.py lines 1064 and 1101); a resume restores the cumulative spend from the ledger, the per-pass records and start time from the earlier metrics, the lease-wait total, and the next pass number from the transcript names (lines 2406 to 2435); the retry-feedback map is loaded only when the mode is on (line 2446); the escape guard’s baseline fingerprint of the three guarded subtrees is taken (parent_repo_fingerprint, line 2455).

Each iteration k from start_pass to max_passes (line 2464) does the following, in this order.

(a) The stopping rule comes first (line 2468): when the cumulative spend already meets the budget, budget_exhausted is set and the loop ends.

(b) The prompt is composed (lines 2472 to 2494): the template, the task package, and an overhead block holding the pass header and, from the second pass or on a resume, the retry notice chosen by compose_retry_notice (lines 925 to 943) for the kind of the previous failure (visible, environment or hidden; the hidden notice names the unmet numbered requirement when unmet_criterion, lines 291 to 316, resolves the failed check through the task’s map). A prompt over 96,000 characters has its overhead trimmed by bound_fail_summary (lines 855 to 866), because the prompt is one command-line argument and the kernel rejects any argument over 128 KiB; run set 017 lost a cell that way. The three injected texts are recorded with their sizes for the ledger.

(c) The agent is invoked under the agent_invocation span (line 2503). On the Anthropic route claude_argv (agent_backends.py line 382) composes the command, containerize_argv (lines 493 to 587) wraps it in docker run with only the workspace mounted, a writable temporary configuration folder, the chosen credential, the lane labels, and a container-id file that is the durable proof the pass ran in a container; invoke_claude_cli (lines 319 to 329) runs it and saves standard output and error verbatim. On a Fabric route run_pass (agent_backends.py lines 829 to 1063) qualifies the model against the endpoint’s live catalogue, acquires a lease (waiting indefinitely while the controller cannot grant one, _acquire_lease lines 768 to 827), starts a heartbeat thread, runs the same seam, and in a finally block (lines 1005 to 1049) stops the container if the exit was abnormal, stops the heartbeat, releases the lease or cancels the admission, and calls back into the runner’s _record_fabric (line 2522), which writes the whole lifecycle into cell_manifest.json under fabric.passes.<k> after every event (line 2531).

(d) Immediately after the agent returns, capture_handoff (lines 1156 to 1193) copies the agent’s HANDOFF.json to handoff.pass<k>.json and validates it, never affecting the verdict. The lease lifecycle’s durations are recorded as timing lines (lines 2564 to 2575), the transcript is split into model time and tool time (timing_lib.transcript_timing, line 445), the turn ledger pass<k>.turns.jsonl is written (line 2589, timing_lib.write_turn_ledger line 1399), and three reported durations are placed inside the agent’s span (lines 2607 to 2634). One line is appended to the run set’s command_log.jsonl with the command (prompt replaced by a placeholder), its size, the back end, the timestamps and the return code (line 2636).

(e) Usage is accounted. On the Claude routes ledger_lib.build_ledger (line 420) attributes every token of the transcript to a category and the rows are appended to token_ledger.jsonl (lines 2649 to 2653); the pass spend is the sum of input, output and cache-creation tokens; the ledger is reconciled against the transcript’s own totals and a gap above half a percent or fifty tokens sets ledger_reconciliation_warning (lines 2656 to 2661); a second model in the transcript and a read of the predecessor’s handoff note are noted (lines 2666 to 2670). On the Codex route the provider’s reported usage is appended to provider_usage.jsonl (line 2673). The cumulative spend advances (line 2679).

(f) The escape guard re-fingerprints the guarded subtrees (line 2690). A changed committed tree is an operator-side deviation: the event is appended to metrics.json’s list and to the execution log (line 2713) and the cell continues. A changed working tree without a commit is a workspace escape: the cell halts as a protocol_violation, invalid (lines 2716 to 2725).

(g) The invocation-error guard reads the transcript’s terminal result line (_invocation_error, lines 1463 to 1505) and, on a Fabric route, the back end’s error text (lines 2731 to 2741). An error halts the cell as agent_invocation_error, invalid, with a per-pass record and a NOT-RUN line (lines 2742 to 2755). A Codex pass with no usage halts as usage_unavailable (lines 2757 to 2772), because missing usage is never recorded as zero.

(h) The visible tests run (line 2782, run_visible_tests lines 809 to 833), inside the container when one is configured (containerize_shell, lines 590 to 604), with a 900-second timeout that counts as a failure. A green container run is confirmed on the host (lines 2814 to 2818); a host failure marks the pass environment_dependent and shows the agent the host’s output (lines 2820 to 2823).

(i) The patch of the workspace against the staged snapshot is written to patches/pass<k>.patch (line 2827, _make_patch lines 1303 to 1323).

(j) When the visible suite passed, the hidden checks run inside the loop (line 2852, _score_hidden_for_pass lines 607 to 638; the write is at 637), writing scoring_outputs/pass<k>.hidden.txt. Scoring here costs no tokens and ends the cell at the pass where the work became correct, which is what makes tokens_to_success mean what its name says (comment at lines 2830 to 2849).

(k) The per-pass record is appended (line 2855) and the loop decides (lines 2869 to 2914): both gates green ends the loop with visible_pass true and the success time stamped; an environment-dependent pass is told the cause once and halted on the second occurrence; a hidden failure behind a green suite sets the hidden notice and records which notice the next pass will receive; a visible failure sets the visible notice with the trimmed test output.

4.11 Finalisation, lines 2917 to 3029

The private Git repository is removed (line 2919), hashes_after.json is written (line 2924), seconds_to_success is computed as the wall clock from the cell’s start to its success minus every second spent waiting for a lease (line 2933), and metrics.json is written at line 3022 with the fields of section 7. The live marker is removed (line 3023), a status still unset becomes completed (line 3026), and _finalize_manifests (lines 3142 to 3228) updates the run set’s run_manifest.json (line 3184) and, under the campaign-wide publish lock, the run index row (lines 3195 to 3223).

5. The state machine of a cell

The states are the values the runner can leave in result["status"] and metrics.json, plus the transient activities named in timings.jsonl. Every transition names its line.

stateDiagram-v2
    [*] --> Resolving : main() 3255, run() 1674
    Resolving --> Exited_without_record : SystemExit 1716..1783, 1953
    Resolving --> Staging : 1959
    Staging --> Refused : missing_start_tree 2034
    Staging --> Injecting : 2079
    Injecting --> Refused : contaminated_task_brief 2092
    Injecting --> Refused : missing_defect_realization 2101/2124
    Injecting --> Refused : missing_feature_realization 2148/2151
    Injecting --> Gating : 2162
    Gating --> Refused : staging_gate_failed 2179
    Gating --> Satisfied_by_start_tree : 2186
    Gating --> Admitted : verdict ok 2190
    Admitted --> Dry_run : 2356
    Admitted --> Pass : 2464
    state Pass {
        [*] --> Stopping_rule : 2468
        Stopping_rule --> Composing : 2472
        Composing --> Invoking : 2503
        Invoking --> Waiting_for_lease : agent_backends 768
        Waiting_for_lease --> Invoking : lease granted
        Invoking --> Capturing : 2554
        Capturing --> Ledgering : 2649/2673
        Ledgering --> Guarding : 2690
        Guarding --> Testing : 2782
        Testing --> Confirming_on_host : 2814
        Confirming_on_host --> Patching : 2827
        Testing --> Patching : 2827
        Patching --> Hidden_scoring : 2852
        Hidden_scoring --> Deciding : 2869
    }
    Pass --> Budget_exhausted : 2469
    Pass --> Protocol_violation : 2718
    Pass --> Agent_invocation_error : 2744
    Pass --> Usage_unavailable : 2760
    Pass --> Succeeded : 2869
    Pass --> Halted_environment : 2878
    Pass --> Pass : retry, next k
    Pass --> Passes_spent : loop end
    Pass --> Crashed_without_record : uncaught error or SIGTERM
    Succeeded --> Finalising : 2919
    Budget_exhausted --> Finalising
    Passes_spent --> Finalising
    Halted_environment --> Finalising
    Protocol_violation --> Finalising
    Agent_invocation_error --> Finalising
    Usage_unavailable --> Finalising
    Finalising --> Completed : 3022, 3026
    Refused --> [*] : _refuse 1859, exit 2
    Satisfied_by_start_tree --> [*] : 1870, exit 0
    Dry_run --> [*] : 3027
    Completed --> [*] : 3027, exit 0
    Exited_without_record --> [*]
    Crashed_without_record --> [*]

Two states in the diagram write no record. Exited_without_record is a SystemExit raised before the cell directory has anything in it, or the folder-reuse refusal that writes nothing on purpose. Crashed_without_record is any uncaught exception or the first stop signal landing inside the pass loop: the signal handler at lines 3230 to 3252 raises SystemExit(143) so that the back end’s finally can release the lease, but the runner itself has no protective block between line 2464 and the metrics write at line 3022, so metrics.json is never written and the scorer later records run_loop_did_not_finish (score_cell.py line 1016). The rulings specification operations/results-integrity/B-rulings-specification.md, section 2.2, records this as the cause of the two undisposed cells of the batches aborted on 2026-09-12.

6. The sequence of one pass on a Fabric route

The participants are the batch driver, the runner, the timer, the back end object, the controller client, the physical controller, the container holding the agent, the scorer module borrowed for hidden checks, and the files. The Anthropic route omits the controller and the lease; everything else is the same.

sequenceDiagram
    participant D as run_batch.py 449
    participant R as run_cell.run() 2464
    participant T as StageTimer (timing_lib 88)
    participant B as DgxClaudeBackend.run_pass (agent_backends 829)
    participant C as FabricClient (dgx_fabric 241)
    participant F as Fabric controller
    participant A as docker run: agent (containerize 437/493)
    participant S as score_cell.run_hidden_scoring 553
    participant X as cell files
    D->>R: run_cell.py --run-set ... (run_batch 179)
    R->>T: open("pass", k) 2466
    R->>R: compose prompt, retry notice 2472..2494
    R->>T: open("agent_invocation") 2503
    R->>B: run_pass(prompt, idempotency_key) 2536
    B->>B: qualify model against catalogue 918 (def 715)
    B->>C: acquire(lane, alias, key) 381
    C->>F: POST /control/v1/admissions
    loop queued or blocked_by_mode (dgx_fabric 408)
        C->>F: GET /control/v1/admissions/id
    end
    C->>F: POST .../claim 416
    F-->>C: lease id, model provenance
    B->>X: cell_manifest.json fabric.passes.k via _record_fabric 2531
    B->>C: HeartbeatTask.start 548 (every min(60, lease/3) s)
    B->>A: invoke_claude_cli(argv) 319
    A-->>F: inference requests through the lease
    A-->>X: transcripts/pass<k>.stream.jsonl, stderr, container_id
    A-->>B: exit status
    B->>C: heartbeat.stop 556, release(outcome) 439
    B-->>R: AgentRunResult (lifecycle, usage, error text) 1051
    R->>T: close("agent_invocation") 2547
    R->>X: handoff.pass<k>.json 2554
    R->>X: pass<k>.turns.jsonl 2589
    R->>T: record lease_wait, admission, agent_process, release 2564..2575
    R->>X: command_log.jsonl 2636
    R->>X: token_ledger.jsonl 2649..2653
    R->>R: parent_repo_fingerprint 2690
    R->>R: _invocation_error 2731
    R->>A: docker run bash -lc visible tests 2782
    A-->>X: scoring_outputs/pass<k>.visible.txt
    R->>R: host confirmation 2818 (pass<k>.visible.host.txt)
    R->>X: patches/pass<k>.patch 2827
    R->>S: _score_hidden_for_pass 2852
    S->>S: checks.py subprocess (score_cell 472)
    S-->>X: hidden_tier_result.json, pass<k>.hidden.txt
    S-->>R: passed, raw
    R->>R: per_pass.append 2855, decide 2869
    R->>T: close("pass") at next iteration 2465 or 2917

7. The records, with their writers and readers

Every file the runner produces, the line that writes it, and the programs that read it. Readers are named by file and line where the reading is a single call, and by file where the reading is spread through the module.

RecordWriterFields or contentReaders
cells/<cell>/cell_manifest.jsonline 2352 (whole record); line 2531 (per-pass fabric.passes.<k> lifecycle, rewritten after every controller event)schema_version, run_id, cell_id, context_window_tokens, model, condition, task_id, phase, seed, max_passes, token_budget, prompt (name, is_lap, variant, artifact_hint, retry_feedback_mode), variant_root, task_dir, start_tree, seeded_defect_injection, feature_removal_injection, staging_gate, visible_test_command, visible_test_commands, documentation_obligation, provider, agent_backend, validation_purpose, invalid_for_primary_analysis, container (image, digest, requirements_lock_sha256, claude_code_version, daemon identity), billing, shell_available, lane, fabric (endpoint, url, lane, requested_alias, priority, lease_seconds, request_preset, passes), createdcheck_cell_manifest_conflict 1584 (the runner itself on reuse); resume 2075 to 2079, 2534; score_cell._load_manifest 1141; score_cell.shell_evidence 887; aggregate_metrics.load_cell 361; aggregate_timings.py 378 (lease lifecycle); capture_evidence.collect_container_proof; gen_run_set_report.py; gen_cell_narrative.py; server.py (live view and reports)
cells/<cell>/hashes_before.json, hashes_after.json2064; 2924 (_sha256_tree 1120){relative path: sha256} of every source file, artefact directories excludedextract_cell_metrics.py; aggregate_timings.py (modification time as a reconstruction anchor)
cells/<cell>/staged_snapshot/2197The post-injection tree, artefacts excluded_make_patch 2827; extract_cell_metrics.py; rescore_tiered.py; regenerate_token_ledgers.py; cleanup_run_set.py (prunes caches under it)
cells/<cell>/workspace/2055 (copy), then the agent through the container mountThe tree the agent editscontainer mount 493 and agent_backends.py 437; score_cell.py (visible, full-suite and hidden steps); capture_evidence.py; the next chained step through run_batch._chain_start_tree 257; cleanup_run_set.py
cells/<cell>/lap_path_manifest_snapshot.json2214lap_files, anchor_files, trace_anchor_sites, probe_comment_sites, payload_metricsaggregate_metrics.load_snapshot 152; extract_cell_metrics.py
<run set>/prompts_snapshot/<template>.md, artifact_hint.v001.<condition>.md2222; 2245The prompt template and the composed hint, verbatimgen_run_set_report.py (prompt provenance); readers of the run set
cells/<cell>/transcripts/pass<k>.stream.jsonl, pass<k>.stderr.txtinvoke_claude_cli 319 to 329 (standard output and error saved verbatim before any parse); on a Fabric route through the same seam from agent_backends.py 979The agent’s whole session as JSON linesledger_lib.build_ledger 420; timing_lib.transcript_timing 445 and turn_ledger 818; _invocation_error 1463; _multi_model 1452; _handoff_note_read 1412; _reported_usage 1401; score_cell.shell_evidence 887, collect_invalid_reasons 976, _transcript_models 1059; aggregate_metrics._pass_transcripts 223; aggregate_timings.py; extract_cell_metrics.parse_transcript; audit_transcripts.py; regenerate_token_ledgers.py; backfill_turn_ledgers.py; server.py 3177 (timings_pass)
cells/<cell>/transcripts/pass<k>.container_idcontainerize_argv 2515 (Docker writes it); agent_backends.py 437 on the Fabric routeThe started container’s identifieragent_backends.stop_container_from_cidfile 516; capture_evidence.collect_container_proof
cells/<cell>/transcripts/pass<k>.turns.jsonl2589 (timing_lib.write_turn_ledger 1399)One line per model request: timestamps, prompt and output tokens, blocks, tool calls, round-trip secondstiming_lib.read_turn_ledger 1434 and pass_timing_rows 1496; aggregate_timings.py; server.py 3177
cells/<cell>/patches/pass<k>.patch2827 to 2828Unified diff of the workspace against the staged snapshot, artefacts removedaggregate_metrics._load_patch 770 and parse_patch 284; gen_cell_narrative.parse_patch; extract_cell_metrics.parse_patch
cells/<cell>/scoring_outputs/staging_gate.txt2175Raw output of the hidden checks on the start treegen_cell_narrative.py
cells/<cell>/scoring_outputs/pass<k>.visible.txt, pass<k>.visible.host.txt2782, 2818 (run_visible_tests 809)Raw test output; a timeout ends with TIMEOUT:score_cell._passes_array 1150 and _agent_visible_outcome 718; gen_cell_narrative.py; the B4 forensics
cells/<cell>/scoring_outputs/pass<k>.hidden.txt637 (inside _score_hidden_for_pass)Raw output of the hidden checks after pass kgen_cell_narrative.py; the B4 forensics
cells/<cell>/hidden_tier_result.json, scoring_outputs/hidden_tier_result.pass<k>.jsonthe task’s hidden_checks/checks.py, through the path the scorer sets in FIVEB_HIDDEN_TIER_OUT (score_cell.py 531 to 534)Which level of evidence answered, per checkgen_run_set_report.py; classify_instrument_errors.py; rescore_tiered.py
cells/<cell>/handoff.pass<k>.jsoncapture_handoff 1184The agent’s HANDOFF.json, copied when it is the agent’s own and not the predecessor’sthe B4 forensics; gen_run_set_report.py
cells/<cell>/token_ledger.jsonl2649 to 2653 (append per pass; rows from ledger_lib.build_ledger 420)Per row: run_id, cell_id, pass_number, event_index, event_type, tool_name, target_path, category, lap_artifact_id, input_tokens, output_tokens, cache_read_tokens, cache_creation_tokens, attribution_method, notesthe runner on resume 2409 to 2415 and _category_totals_from_ledger 3057; aggregate_metrics.load_cell 366 (its absence refuses the cell at 1113); regenerate_token_ledgers.py 95; server.py 1552 and 1745
cells/<cell>/provider_usage.jsonl2673Codex usage per passthe runner on resume 2416 to 2421; no campaign table, because the aggregator requires a token ledger
<run set>/command_log.jsonl2636run_id, cell_id, pass, argv (prompt replaced), prompt_chars, agent_backend, started, ended, returncodeaggregate_timings._command_log 260; capture_evidence.py 119
<run set>/execution_log.md_log_not_run 3135, _log_reached 3122, _log_parent_repository_change 796; also run_batch._log_not_run 227 and server.py 2369One Markdown line per event: NOT-RUN, REACHED, DEVIATIONserver.py 2856 (_not_run_entries); gen_run_set_report.py; lint_report_prose.py
cells/<cell>/timings.jsonlStageTimer._append (timing_lib.py 118) through stage 147, open 187, close 205 and record 229Per line: schema, run_id, cell_id, stage, phase, pass, parent, depth, started_at, ended_at, seconds, ok, measured, attrs_timing_summary 3032 (into metrics.stage_timings); score_cell.py 1277 (appends its own lines) and _refresh_stage_timings 1403; aggregate_timings.py 499; server.py 901
cells/<cell>/timing_current.jsonStageTimer._mark_live (timing_lib.py 125); removed by clear_live 270The innermost activity in progressserver.py live view
cells/<cell>/metrics.json3022 (a completed, halted or exhausted cell); 1932 (a reached step)cell_id, run_id, project_id, condition, task_id, agent_backend, provider, validation_purpose, phase_id, seed, lap_profile_id, model_class, model, context_window_tokens, fabric_endpoint, model_provenance, task_dir, pocket, started_at, completed_at, seconds_to_success, lease_wait_seconds, passes_run, budget_exhausted, token_budget, cumulative_spend, visible_pass, agent_invocation_error, ledger_reconciliation_warning, invalid_for_primary_analysis, not_run_reason, environment_dependent_code, prior_invalidity_carried_forward, parent_repository_changed_by_commit, predecessor_handoff_present, predecessor_handoff_read, prompt_variant, artifact_hint_applied, retry_feedback_available, retry_feedback_mode, per_pass (each with pass, spend, cumulative, visible_pass, hidden_pass, environment_dependent_code, returncode, handoff_present, handoff_valid, handoff_problems, injections, category_tokens, and retry_feedback on a hidden failure), category_totals, stage_timingsthe runner on resume 1848, 2423; run_batch._invocation_errored 464, _reached_without_an_agent 475, _cell_was_halted 327, _cell_is_terminal 365; score_cell.collect_invalid_reasons 976 (existence), _run_halt_reasons 1094, _refresh_stage_timings 1403 (rewrites stage_timings); aggregate_metrics.process_run_set 1098 and load_cell 361; aggregate_timings.py; cleanup_run_set._finished_without_a_score 65; gen_run_set_report.py; gen_cell_narrative.py; server.py; rewritten by regenerate_token_ledgers.py and repair_invocation_errors.py
cells/<cell>/scoring_summary.json and scoring_outputs/scoring_summary.json (reached step only)1928 to 1931The reached-step verdict with score_states.correctness_gate equal to satisfied_by_start_treescore_cell.reached_step_summary 1236 (returns it untouched unless forced); run_batch._note_chain_outcome 296; aggregate_metrics.reached_cell_metrics 1046
<run set>/run_manifest.json_finalize_manifests 3184schema_version, run_id, project_id, conditions, task_ids, model, agent_backend, fabric_endpoints, cost_route, external_model_charge_usd, validation_purpose, max_passes, prompt_versions, cells (per cell: status, not_run_reason, visible_pass, budget_exhausted, fabric_endpoint, invalid_for_primary_analysis)gen_run_set_report.py; gen_iteration_summary.py; cleanup_run_set._max_passes 55; attest_batch_outcome.py; restore_reached_step.py; server.py
5. Run Sets/run_index.csv_finalize_manifests 3195 to 3223, one row per run set under the publish lock; the status column is later closed by run_batch.set_run_index_status 773run_id, date, status, project_id, task_ids, conditions, lap_profiles, model_class, primary_read, valid_for_primary_analysisupdate_monitoring.py; server.py (runsets); publish_snapshot.py

The lineage from these files to the campaign tables and the published number is drawn below. The closeout chain that performs each arrow is run_batch.py lines 564 to 596 and 852 to 897 (chapter 10).

flowchart LR
    subgraph cell [cells/<cell>/ written by run_cell.py]
        M[metrics.json 3022]
        L[token_ledger.jsonl 2649]
        TR[transcripts/pass k .stream.jsonl 319]
        TL[transcripts/pass k .turns.jsonl 2589]
        P[patches/pass k .patch 2827]
        CM[cell_manifest.json 2352]
        TI[timings.jsonl timing_lib 118]
        SN[lap_path_manifest_snapshot.json 2214]
        VO[scoring_outputs/pass k .visible.txt 2782]
    end
    subgraph runset [run set level]
        CL[command_log.jsonl 2636]
        EL[execution_log.md 3122 3135 796]
        RM[run_manifest.json 3184]
    end
    SC[score_cell.py 1258 writes scoring_summary.json 1394]
    VO --> SC
    M --> SC
    TR --> SC
    CM --> SC
    TI --> SC
    AG[aggregate_metrics.py process_run_set 1098 compute_cell 398]
    M --> AG
    L --> AG
    TR --> AG
    P --> AG
    SN --> AG
    SC --> AG
    AT[aggregate_timings.py collect 1054]
    TI --> AT
    CL --> AT
    CM --> AT
    TL --> AT
    CFM[(6. Metrics/cell_factor_matrix.csv)]
    PRM[(6. Metrics/paired_response_matrix.csv)]
    PCM[(6. Metrics/pass_count_matrix.csv)]
    TM[(6. Metrics/time_matrix.csv)]
    TSI[(6. Metrics/token_surface_inputs.csv)]
    CT[(6. Metrics/cell_timings.csv stage_timings.csv pass_timing.csv timing_summary.json)]
    AG --> CFM
    AG --> PRM
    AG --> PCM
    AG --> TM
    AG --> TSI
    AT --> CT
    UM[update_monitoring.py 70]
    CFM --> UM
    UM --> MON[(7. Monitoring/*)]
    RP[gen_run_set_report.py write_report]
    CFM --> RP
    M --> RP
    SC --> RP
    RM --> RP
    EL --> RP
    RP --> RES[(run set results/results.md and 8. Reports/Run Set Results)]
    SV[server.py cells coverage timings]
    CFM --> SV
    CT --> SV
    SV --> UI[the results interface]
    PS[publish_snapshot.py build 350]
    SV --> PS
    PS --> PORTAL[(the public portal snapshot)]
    PE[publish_experiment_results.py publish 221]
    RES --> PE
    PE --> PORTAL

8. The loops and the waits

The pass loop (lines 2464 to 2914) runs at most max_passes iterations and ends early on success, on a halt, or when the cumulative spend meets the budget; the budget is checked at the top of each iteration, so the last pass may overrun it. Inside one pass, on a Fabric route, _acquire_lease (agent_backends.py lines 768 to 827) retries a retryable admission failure indefinitely at FIVEB_LEASE_RETRY_SECONDS intervals (default 60 seconds), recording each round as a waiting_for_lease event and the seconds waited in lease_wait_seconds, which the runner subtracts from the time to success; a non-retryable failure (bad token, unknown alias) raises at once. FabricClient.acquire (dgx_fabric.py lines 381 to 430; the polling branch at 408) polls a queued or mode-blocked admission every poll_seconds plus jitter until it is ready, claims it, or recovers a saved lease. The heartbeat thread (dgx_fabric.py lines 525 to 563) renews the lease every min(60, lease_seconds / 3) seconds until stopped, and a failure inside it is raised only when it is stopped. The visible tests and the hidden checks each have a 900-second limit (VISIBLE_TEST_TIMEOUT_S line 332; score_cell.TEST_COMMAND_TIMEOUT_S). The escape guard’s git status call has a 300-second limit (line 736). No other loop in the runner waits on anything external.

9. The guards and refusals

GuardWhereWhat it refuses or records
Back-end and route validity1675, 1713 to 1752Unknown back end; Fabric back end without a container, endpoint or allowed lane; Anthropic credential on the dgx_claude route; an Anthropic route whose billing cannot authenticate
Folder reuse1953 (check_cell_manifest_conflict 1584)Any change of model, prompt variant, retry-feedback mode or Fabric endpoint against the folder’s earlier manifest; writes nothing
Brief contamination2089 (lint_task_briefs.scan_task_dir)A task brief that names the study; the cell is refused before any spend
Hint contamination2235 to 2242A composed artifact hint that trips the same lint, except the pattern that forbids naming documentation artefacts
Staging gate2172 (run_staging_gate 961)A start tree that already passes its hidden checks, or a task with no executable scorer
Sticky invalidity1845 to 1857A resumed cell inherits an earlier invalid verdict
Prompt length2486Trims the retry notice so the argument stays under the kernel’s limit
Workspace escape2690 to 2725 (parent_repo_fingerprint 675)A working-tree change under 1. Harness, 2. Project Library or 4. Task Library halts the cell; a committed change is recorded as a deviation
Invocation error2731 to 2755 (_invocation_error 1463)A pass that ended on an API error or a dead stream is an environment failure, never a capability measurement
Usage presence2757 to 2772A Codex pass without reported usage is not recorded as zero
Host confirmation2814 to 2823A container-green suite must also pass on the host before it can end the cell
Patch size_make_patch 1319 to 1323A filtered patch above 50 MiB raises
Signal handling3230 to 3252Only the first stop signal unwinds; later ones are dropped so the lease release can complete

10. Every unhappy path, with its trigger and its record

TriggerCode pathWhat is writtenStatus and exitDownstream effect
Unknown back-end nameagent_backends.select_backend 373, called at 1675nothingSystemExit, exit 1The driver sees a non-zero return and no metrics.json; the cell has no disposition (B rulings specification, section 5)
Fabric back end without FIVEB_CELL_CONTAINER1714 to 1716nothingSystemExitas above
Endpoint, lane or credential misconfiguration1719 to 1749nothingSystemExitas above
Billing route cannot authenticate1752 (assert_billing_ready 427)nothingSystemExitas above
Task folder or variant root missing1777, 1782nothingSystemExitas above
Folder reuse conflict1953 to 1958nothing (live marker cleared)SystemExitas above
Missing or empty start tree of a chained step2034execution_log.md NOT-RUN, run_manifest.json cell status not_run, run index rownot_run, exit 2The driver’s chain bookkeeping refuses the rest of the chain (run_batch._note_chain_outcome 296)
Contaminated task brief2092same as above, reason contaminated_task_brief:<file>:<line>:<label>not_run, exit 2The task must be repaired before any cell on it can run
Missing defect realisation or a diff that does not apply2101, 2124same, reason missing_defect_realization:<condition>not_run, exit 2The realisation for that build must be authored
Missing feature realisation2148, 2151same, reason missing_feature_realization:<condition>not_run, exit 2materialize_feature_bases.py must be run for the task
Staging gate refuses2179scoring_outputs/staging_gate.txt, then the not_run records, reason staging_gate_failed:hidden_scorer_missing or staging_gate_failed:start_tree_passes_hiddennot_run, exit 2A missing scorer is an instrument fault (lint_hidden_scorers.py); a passing start tree is a staging defect
Staging gate satisfied on a chained step2186scoring_summary.json (twice), metrics.json with zero counts, execution_log.md REACHED, run manifestssatisfied_by_start_tree, exit 0The driver does not score the cell (run_batch 475 to 490) and the chain continues from the same tree; the aggregator builds its row at zero cost (aggregate_metrics.reached_cell_metrics 1046)
Contaminated artifact hint2240the not_run records, reason contaminated_artifact_hintnot_run, exit 2The hint template or the inventory must be corrected
Dry run2356 to 2360cell_manifest.json, run manifestsdry_run, exit 0No passes, no scoring
Budget met at the top of a pass2468 to 2470normal finalisation; budget_exhausted truecompleted, exit 0The scorer scores the last tree; pass_count_matrix.csv records the exhaustion
Prompt longer than 96,000 characters2486 to 2489the trimmed overhead is what the agent sees; the injection sizes record the trimmed lengthcontinuesNone beyond a shorter notice
Lease cannot be granted (maintenance, reboot, no capacity)agent_backends._acquire_lease 768 to 827cell_manifest.json fabric.passes.<k>.events gains waiting_for_lease rows; lease_wait_seconds accumulateswaits indefinitelyThe live view says the cell is waiting; the wait is subtracted from seconds_to_success
Non-retryable admission failuresame, re-raised; caught by run_pass’s broad handler 1001 to 1004pass<k>.stderr.txt holds the redacted error; the transcript is created emptyback-end error_text set, so the invocation-error guard halts the cell (2738)see the invocation-error row
Agent process exits non-zero, or the transcript’s last result line is an API error or an aborted stream2731 to 2755 (_invocation_error 1463)per_pass record with agent_invocation_error, execution_log.md NOT-RUN, then metrics.json with agent_invocation_error true and invalid_for_primary_analysis trueagent_invocation_error, exit 0The driver does not score (run_batch 464 to 468); two in a row abort the batch (run_batch 1034 to 1041); the aggregator refuses the cell for want of a scoring record (1121)
Codex pass reports no usage2757 to 2772as above with usage_unavailableusage_unavailable, exit 0as above
Guarded working tree changed during a pass2716 to 2725metrics.json with not_run_reason workspace_escape_detected..., invalidprotocol_violation, exit 0The scorer adds run_halted:workspace_escape_detected (_run_halt_reasons 1094); operator review required
Guarded committed tree changed during a pass2698 to 2715execution_log.md DEVIATION, metrics.json parent_repository_changed_by_commitcontinuesRecorded, not judged; the report’s deviations section prints it
Visible test command exceeds 900 secondsrun_visible_tests 825 to 831pass<k>.visible.txt ending in TIMEOUT:, return code 124counts as a visible failureThe next pass receives the visible notice with the truncated log
Container-green suite fails on the host2814 to 2823pass<k>.visible.host.txt; environment_dependent_code on the pass and the cellfirst time: the environment notice; second consecutive time: halt at 2878 to 2884 with not_run_reason and a NOT-RUN lineThe halt sets no invalid flag, so _run_halt_reasons (which returns nothing unless invalid_for_primary_analysis is true, score_cell.py line 1110) ignores it and the cell is scored on the tree it left, valid unless another reason applies. This is what the code does; whether it is what the operator intends is recorded as an open question in section 14
Filtered patch above 50 MiB_make_patch 1319 to 1323nothing further; the exception is uncaughtcrash, no metrics.jsonThe scorer records run_loop_did_not_finish (score_cell.py 1016); the aggregator refuses the cell for want of a ledger row or verdict
Hidden checks raise instead of answering_score_hidden_for_pass 629 to 633pass<k>.hidden.txt is not written; hidden_pass is Nonethe visible verdict alone ends the loop (2869)The scorer’s own hidden step records hidden_scorer_missing or the real answer
Ledger disagrees with the transcript by more than half a percent or fifty tokens2656 to 2661ledger_reconciliation_warning true in metrics.jsoncontinuesupdate_monitoring.py counts the warnings; the aggregator’s reconcile_check (244) raises on a larger mismatch
Timing line cannot be writtentiming_lib._append 118 to 123the line is lostcontinuesstage_timings is incomplete; aggregate_timings.py reconstructs from file times where it can
First stop signal during a pass3230 to 3252, then agent_backends.run_pass 1005 to 1049the back end stops the container, releases the lease, and records the lifecycle in cell_manifest.json; metrics.json is never written; timing_current.json may remainSystemExit(143)The scorer, if the driver runs it, records run_loop_did_not_finish; otherwise the cell has no disposition (B rulings specification, sections 2.2 and 5)
run_index.csv absent3188run_manifest.json onlycontinuesThe campaign index lacks the row until a later cell writes it

11. The metrics this chapter produces and where each goes

metrics.json is the runner’s contract with the aggregator (aggregate_metrics.load_cell 361, compute_cell 398). cumulative_spend, per_pass[].spend and category_totals feed the token columns of cell_factor_matrix.csv, of which tokens_to_success is the primary metric (run_defaults.json, key primary_metric); the aggregator recomputes them from the ledger rows rather than copying the runner’s sums. passes_run and budget_exhausted feed pass_count_matrix.csv and the pass-at-budget file. seconds_to_success and lease_wait_seconds feed time_matrix.csv. visible_pass on the runner’s record is advisory; the verdict of record is the scorer’s scoring_summary.json, and the B3 verification found that the aggregator’s visible column reads the scorer’s per-pass parse rather than the runner’s flag (tracker item 2.44). invalid_for_primary_analysis, not_run_reason and validation_purpose are read by the scorer’s collect_invalid_reasons (976) and become invalid_reasons on the summary, which the aggregator copies unchanged into the matrix (aggregate_metrics.py 505, per the B2 count-provenance memo). prompt_variant, artifact_hint_applied, retry_feedback_mode, retry_feedback_available, predecessor_handoff_present, predecessor_handoff_read, model_provenance and fabric_endpoint are analysis dimensions carried into the matrix as columns. stage_timings is the per-cell roll-up of timings.jsonl; the campaign timing tables are built from the line-by-line file by aggregate_timings.py, not from the roll-up.

12. The tests that exercise the runner

The harness test suite under 5. Experiment/1. Harness/scripts/tests/ (54 entries) drives the runner through its two seams, invoke_claude_cli (319) and run_visible_tests (809), which are replaced by fakes; the real agent is never called. The files that name the runner are test_run_cell.py, test_integration_smoke.py, test_isolation_contract.py, test_prompt_variant_and_retry_feedback.py, test_chained_gate_satisfied.py, test_start_tree_staging_gate.py, test_reached_step_scoring.py, test_admission_identity.py, test_lease_wait.py, test_fabric_endpoint_wiring.py, test_agent_backends.py, test_run_batch.py, test_run_batch_sequences.py, test_resume_into.py, test_stop_and_evidence.py, test_closeout_progress.py, test_lane_coordinator.py, test_timing_lib.py, test_score_cell.py, test_run_set_report.py and conftest.py. Which unhappy path of section 10 each test covers was not mapped in this pass; the maintenance contract of the specification (section 7) makes that mapping a column of the dependency table.

13. The dated incidents that shaped the code

The comments in the file carry the history that explains its shape, and a reader who does not know it will mistake a guard for a mistake. The escape guard exists because on 2026-07-28 an agent reached the frozen documented build through an absolute path and committed to the parent repository (line 678); it was narrowed to three subtrees on 2026-08-15 after plan-document commits from a companion session halted four attempts in run set 009 (line 688). The retry notices were rewritten on 2026-08-15 because the earlier wording told the agent a hidden oracle existed, and one agent went looking for it and found the hidden scoring assets on 2026-07-26 (lines 181 to 197). The staging gate dates from 2026-08-08, when three feature cells of run set 007 were found to have been staged with the feature already built (line 965). The invocation-error guard dates from 2026-08-09, when a credit balance ran out mid-batch and every later pass scored as an agent failure (line 1467). Hidden scoring moved inside the loop on 2026-08-16 because run set 009 had recorded 68 percent of its tokens after the work was already right (lines 2830 to 2849). The host confirmation dates from the same day, when a cell wrote a container path into production code and was failed by a run it never saw (lines 2785 to 2812). The environment notice dates from 2026-08-17 (line 206). The billing route became explicit after run set 016 ran on the metered account by accident (lines 380 to 389). The signal handler’s repeat guard dates from the stops of run sets 038 and 039 on 2026-08-29 (lines 3233 to 3249). Sticky invalidity dates from 2026-08-15 (line 1838). The single prompt is v006 since 2026-09-03 (lines 81 to 103). The artifact hint and the criterion notice are amendments A8 and A9 of 2026-09-01 (lines 110 to 131, 235 to 253). The chained-step ruling is Decision Sheet item 54 of 2026-09-03 (line 953), and the carried handoff note is amendment A12 of the same day (line 2046).

14. The weakest claim what was not checked and the token line

The weakest claim is the reader lists in the record table of section 7. The writers are read from the runner’s own lines and are exact. The readers were gathered by name search across the harness and interface folders, and a reader that opens a file by a constructed path the search did not match would be missing; the interface’s server.py, at 211,295 bytes, was searched and not read, so its reader entries name the file without a line except where the search returned one. The second weakest claim is the treatment of the environment-dependent halt in section 10: the code at lines 2878 to 2884 sets a reason without the invalid flag, and the scorer’s halt reader at line 1110 requires the flag, so the halt leaves a valid, scored cell; that reading is of the code alone, and no cell in the campaign was checked to confirm that a cell halted this way carries a valid verdict on disk.

Not checked: which test file covers which exit path (section 12); whether score_cell.py line 1277, where the scorer appends to the same timings.jsonl, is the only place another program writes into a cell’s timing file; the exact line at which agent_backends.py line 981 passes through invoke_claude_cli on the Fabric route (the runner injects itself as runner=invoke_claude_cli at line 2382, and the back end calls self.runner at line 981); whether gen_run_set_report.py reads handoff.pass<k>.json or only the metrics fields about it; and the Codex route’s provider_usage.jsonl was not traced beyond the runner because no valid cell has run on it since run set 030.

Token line: the session that wrote this chapter and the specification beside it had consumed 453,911 tokens of context by the time writing began, measured as the difference of the remaining-token counter (15,000,000 at the start, 14,546,089 at the last check before writing), almost all of it the digest and the source ranges read for verification; the two documents add their own length on top, and the job ledger carries the cost. No harness script was run and no model tokens were spent by the harness.