This chapter was reviewed against the owning scripts at the verified commit. The flow model is used only for the chapter 02 subgraph. Its nodes and edges are named beneath the diagrams so that the rendered account can be checked against the model rather than mistaken for a new execution trace.

1. The purpose and the position in the life of the experiment

A batch is a planned group of measured experiment units. A cell is one unit defined by a condition, task, seed, and optional phase. A run set is the unique directory that holds the records and workspaces for one execution of a batch. The batch driver turns the batch specification into that run set, runs each cell in order, and closes the run even when execution stops early. The lane coordinator is the program that runs independent batch drivers at the same time for named hardware endpoints, while keeping each child process and its environment separate.

The ordinary invocation begins with 5. Experiment/10. User Interface/launch_batch.sh, which prepares the narrow execution environment, records the batch specification in version control, and executes run_batch.py. A user or scheduler may also call the Python entry point directly. The concurrent route begins with 5. Experiment/10. User Interface/launch_lanes.sh, which records every batch named by a lane plan and executes lane_coordinator.py. The coordinator invokes one run_batch.py child per lane. It does not merge cell execution into one process, because run_batch.py mutates its process environment for the selected Fabric endpoint.

Before a cell is staged, preflight checks whether the task declarations, model choice, cost ceiling, credentials, and selected machine are usable. During the batch, the driver invokes run_cell.py and score_cell.py, carries a chained task forward from its predecessor’s workspace, and may resume a hidden failure when passes and budget remain. At the end it writes the batch outcome and closeout progress, updates the active batch and run index, and invokes the closeout programs that feed campaign tables, monitoring, reports, cleanup, and publication. The coordinator additionally writes lane provenance and returns one JSON account of all lane outcomes.

The mechanism therefore sits between an approved specification and the campaign’s derived records. Its outputs are not a single result value. They are the run set folder, cell records and execution log, batch_outcome.json, closeout_progress.json, status surfaces, and, for concurrent execution, lane_provenance.json and the coordinator’s JSON result. The scorer reads the cell records, the aggregator reads the cell and scoring records, and the interface reads the status and provenance records.

2. The reader’s map of the owning files

2.1 The batch driver

run_batch.py is the large orchestrator, 53,836 bytes. Its own numbered comments divide the work into a thin command-line wrapper, a per-cell driver, and a closeout chain. Lines 1 to 62 contain the module contract, imports, and the hidden-failure rule. Lines 64 to 142 contain base-path resolution, signal forwarding, and default_invoke, the subprocess seam. Lines 145 to 215 allocate a run set and construct child arguments. Lines 218 to 340 read JSON, resolve chained start trees, and log or classify chain outcomes. Lines 449 to 540 are _run_one_cell, the main per-cell flow. Lines 543 to 595 write the active marker and define the closeout stages. Lines 601 to 770 construct the stop reason, batch outcome, closeout progress, and captured child output. Lines 773 to 831 update the run index under a lock. Lines 833 to 895 are closeout and its stage finalisation.

The main flow is run_batch in 5. Experiment/1. Harness/scripts/run_batch.py lines 913 to 1068. It validates duplicate cell identities and the single-model rule, applies specification environment values, creates or resumes the run set, iterates cells, applies the circuit breaker, and enters closeout from a finally block. The command entry point is main in 5. Experiment/1. Harness/scripts/run_batch.py lines 1071 to 1113. The file’s finalisation helpers are _write_batch_outcome lines 617 to 650, _write_closeout_progress lines 685 to 718, set_run_index_status lines 773 to 831, _closeout lines 833 to 850, and _closeout_locked lines 852 to 895.

2.2 The lane coordinator

lane_coordinator.py is 24,904 bytes. Lines 1 to 76 contain the concurrency contract, imports, labels, and the environment-key list. Lines 78 to 105 define the LaneSpec record and lines 179 to 193 define LaneResult. The environment and container helpers are build_lane_env in 5. Experiment/1. Harness/scripts/lane_coordinator.py lines 107 to 133, lane_container_label lines 136 to 150, and stop_lane_containers lines 153 to 176. The child launch and validation region is default_lane_invoke lines 199 to 223 and _validate_lane_batch_endpoint lines 225 to 257.

The coordinator main flow is run_lanes in 5. Experiment/1. Harness/scripts/lane_coordinator.py lines 259 to 387. It refuses duplicate lane identities and destinations, starts all children, polls them independently, and constructs the coordinator record. Cancellation is cancel_lane lines 390 to 403. Finalisation is write_lane_provenance lines 406 to 422 and publish_lane_results lines 425 to 445. Plan loading is load_lane_specs lines 448 to 493, and the command entry point is main lines 496 to 519.

2.3 The preflight and sequence files

preflight_batch.py is 32,120 bytes. Its cost-model loader is load_cost_model in 5. Experiment/1. Harness/scripts/preflight_batch.py lines 97 to 105. The ordinary credential gate is probe_spendability lines 107 to 154, cost estimation is estimate_cost lines 157 to 171, and the host and container Fabric probes are probe_dgx_host lines 174 to 290 and probe_dgx_container lines 292 to 345. Task registration and sequence validation are probe_task_registrations lines 348 to 445. The file-only dispatcher is preflight_static lines 447 to 524, the Fabric dispatcher is preflight_dgx lines 527 to 568, the overall dispatcher is preflight lines 571 to 608, and the command entry point is main lines 626 onward.

task_sequences.py is 17,628 bytes. load_chains in 5. Experiment/1. Harness/scripts/task_sequences.py lines 127 to 192 reads task manifests and constructs predecessor and chain maps. validate_sequence_order lines 222 to 289 checks gaps, order, and intervening cells. order_cells in 5. Experiment/1. Harness/scripts/task_sequences.py lines 292 to 326 can produce an ordered permutation, while prune_orphan_steps in the same file is verified at lines 327 to 358. The driver uses the loader and the validation result; the batch loop itself decides whether a failed predecessor breaks the chain.

2.4 The lock, shell launchers, and attestation

publish_lock.py is the small shared-publication lock. _lock_path in 5. Experiment/1. Harness/scripts/publish_lock.py lines 31 to 34 creates the lock directory and names the sentinel. publish_lock lines 37 to 55 opens it, blocks on an exclusive flock, and releases it in a finally block.

run_tranche.sh is the ordered shell entry point. Its numbered comments define the preflight, dependency refresh, image selection, detached launch, and monitoring notice at lines 6 to 15. The actual checks begin at lines 36 to 107 and the source-variant preparation begins at lines 108 to 140. launch_batch.sh reads and exports the execution environment at lines 9 to 30, commits the named batch specification at lines 56 to 107, and executes the driver at lines 110 to 111. launch_lanes.sh validates its plan at lines 9 to 14, loads shared environment values at lines 16 to 27, commits each lane batch under a launch lock at lines 29 to 49, and executes the coordinator at lines 51 to 53.

attest_batch_outcome.py is the repair path for older or silent run sets. The digest verifies count_from_files at lines 61 to 90, build_record at lines 91 to 138, write_record at lines 139 to 143, and main at lines 144 to 185. It derives counts from file existence, keeps human context under an attestation field, writes batch_outcome.json, and attempts the run-index update.

3. The inputs

3.1 Command-line arguments

The Python batch driver parses its arguments in main in 5. Experiment/1. Harness/scripts/run_batch.py lines 1071 to 1084. The coordinator parses its arguments in main in 5. Experiment/1. Harness/scripts/lane_coordinator.py lines 496 to 506. The preflight command parses its arguments in main in 5. Experiment/1. Harness/scripts/preflight_batch.py lines 641 to 647. The verified arguments are:

Program and argumentParser lineDecision made
run_batch.py --batch1073Selects the batch specification JSON to load.
run_batch.py --skip-preflight1074Bypasses the preflight call when present.
run_batch.py --max-usd1076Supplies the worst-case cost ceiling to preflight.
run_batch.py --resume-into1078Selects an existing run set whose cells can be resumed or skipped.
run_batch.py --run-set-destination1082Supplies the not-yet-existing destination for a new run set.
lane_coordinator.py --plan500Selects the explicit lane-plan JSON.
lane_coordinator.py --base-dir502Overrides the experiment root used to resolve relative paths.
lane_coordinator.py --provenance504Selects an optional coordinator-level provenance file.
preflight_batch.py --batch626 onwardSelects the specification for direct preflight.
preflight_batch.py --max-usd626 onwardSupplies a direct cost ceiling.
preflight_batch.py --json626 onwardRequests machine-readable rendering.

The shell launchers add their own positional contract before these parsers run. run_tranche.sh requires exactly one batch path at lines 36 to 52. launch_lanes.sh requires its first positional value to be a plan path at lines 9 to 14, then passes later arguments to the coordinator. launch_batch.sh passes all arguments through to run_batch.py at line 111.

3.2 Environment variables

VariableRead or set atDecision made
FIVEB_FABRIC_ENDPOINTrun_batch.py lines 430 to 446; selected by preflight lines 541 to 544Selects the named Fabric endpoint when the specification does not provide one.
FIVEB_CELL_CONTAINERrun_batch.py lines 440 to 446 and preflight_batch.py lines 310 to 312Selects the measured-cell image and decides whether the container probe can run.
DGX_SPARK_FABRIC_URL, DGX_SPARK_FABRIC_TOKENpreflight_batch.py lines 194 to 200 and shell setup lines 26 to 27Supplies the DGX controller address and authentication for the host probe.
IGX_THOR_FABRIC_URL, IGX_THOR_FABRIC_TOKENdgx_fabric.py endpoint resolution, forwarded by launch_lanes.sh lines 22 to 27Supplies the IGX Thor controller pair when that endpoint is selected.
DGX_SPARK_FABRIC_LANEpreflight_batch.py lines 188 to 204Chooses the authorized maintenance lane.
DGX_SPARK_FABRIC_MODELpreflight_batch.py lines 188 to 190 and 566 to 568Supplies the Fabric model alias when the specification does not.
DGX_SPARK_FABRIC_PRIORITYpreflight_batch.py lines 249 to 254Sets admission priority, defaulting to 50.
DGX_SPARK_FABRIC_LEASE_SECONDSpreflight_batch.py lines 249 to 266Sets the probe lease duration, defaulting to 900 seconds.
FIVEB_LANE_OWNER, FIVEB_LANE_ENDPOINT, FIVEB_LANE_RUN_IDlane_coordinator.py lines 121 to 133Identify the child lane for logging, labels, and cleanup.
ANTHROPIC_API_KEY, FIVEB_CELL_CREDENTIALS, FIVEB_BILLINGlaunch_batch.sh lines 24 to 29 and 74 to 77Are cleared for local Fabric work and selected for the ordinary Claude route.

3.3 Files read

File or pathReader and lineDecision made
Batch specification JSONrun_batch.py lines 1086 to 1087Supplies project, cells, model, chain, budget, and Fabric fields.
Lane plan JSONlane_coordinator.py lines 458 to 493Supplies one explicit LaneSpec per child.
Per-lane batch specificationlane_coordinator.py lines 234 to 256Is checked for endpoint and model agreement before launch.
config/run_defaults.jsonpreflight_batch.py lines 97 to 105Supplies the cost model, with verified fallback constants if unavailable.
4. Task Library/*/*/task_manifest.jsontask_sequences.py lines 127 to 145 and preflight_batch.py lines 367 to 404Supplies predecessor declarations and measured-run validity.
Run-set template 5. Run Sets/001-*run_batch.py lines 156 to 171Supplies the initial directory tree copied for a new run set.
Existing run set and its cells/ directoryrun_batch.py lines 940 to 953Determines whether an in-place resume is valid.
Cell metrics.json and scoring_summary.jsonrun_batch.py lines 464 to 539Determines invocation error, reached state, terminal state, hidden result, budget, and resume eligibility.
run_index.csvrun_batch.py lines 799 to 827Supplies the row whose status is updated at closeout.
Shell environment file .envlaunch_batch.sh lines 26 to 29 and launch_lanes.sh lines 22 to 27Supplies controller values before Python starts.

3.4 Network endpoints and subprocess boundaries

The owning files do not use a network client for ordinary batch orchestration. Their external boundary is the subprocess interface, with the preflight client boundary recorded below.

Endpoint or external boundaryClient function and lineDecision or effect
Sibling harness programsrun_batch.py default_invoke, lines 113 to 142Starts cell and closeout subprocesses and captures closeout output when required.
A lane batch driverlane_coordinator.py default_lane_invoke, lines 199 to 223Starts one run_batch.py child with its copied environment.
Fabric controller status, catalogue, admission, lease, and probe routespreflight_batch.py probe_dgx_host, lines 207 to 287, through FabricClientChecks maintenance, lane capacity, alias compatibility, lease admission, heartbeat, inference, and release.
Measured-cell bridge networkpreflight_batch.py probe_dgx_container, lines 316 to 345, through Docker and dgx_connectivity_probe.pyConfirms that the container can use the selected controller route.
Docker container cleanuplane_coordinator.py stop_lane_containers, lines 165 to 176Lists and stops only containers carrying the lane-scoped label.

The endpoint client, rather than the batch driver, owns controller routes and inference calls. The batch and lane coordinators own the subprocess boundary and the records that explain what those calls produced.

4. The happy path in order

4.1 The launch preparation

The shell launchers establish the reproducible boundary before Python begins. launch_batch.sh in 5. Experiment/10. User Interface/launch_batch.sh lines 14 to 30 narrows PATH, clears credentials that local cells must not receive, selects the cell image, and reads the Fabric values. Lines 56 to 107 find the batch path, commit that specification before launch, and make a failed local commit a refusal to start. A failed push is advisory because the local commit still ties the run to its plan. launch_lanes.sh performs the corresponding plan check and shared environment setup at lines 9 to 27, then serializes Git staging and commits for all named batch files at lines 29 to 49 before it executes the coordinator.

The reason given by the shell comments is evidential rather than cosmetic. A measured result must be matched to the exact plan, and the browser request must not wait for a slow repository operation. In the lane case, the short Git lock prevents two background launchers from colliding on the repository index before either reaches the Python coordinator.

4.2 The preflight gate

main in 5. Experiment/1. Harness/scripts/run_batch.py lines 1071 to 1113 loads the specification and calls preflight unless --skip-preflight is present. The dispatcher preflight in 5. Experiment/1. Harness/scripts/preflight_batch.py lines 571 to 608 chooses the Fabric route for dgx_codex and dgx_claude; otherwise it runs the spendability and cost gates. preflight_static lines 447 to 524 checks task registrations, sequence order, cell folder identities, and the one-model constraint.

For a Fabric batch, preflight_dgx lines 527 to 568 resolves the endpoint once, runs probe_task_registrations, then the host probe and, only if that succeeds, the container probe. probe_dgx_host lines 174 to 290 checks the controller state and lane capacity, confirms the model alias and backend compatibility, acquires a probe lease, sends a heartbeat, makes the protocol probe, and releases it. probe_dgx_container lines 292 to 345 exercises the same route from the measured-cell bridge network. For the ordinary route, probe_spendability lines 107 to 154 makes the minimal call through the same invocation seam as a cell and rejects missing output or an invocation error. estimate_cost lines 157 to 171 calculates the worst case, and preflight compares it with the supplied ceiling.

A refusal is printed as a preflight result and main returns exit code 1 at lines 1100 to 1105. No run set is created. This ordering makes the expensive and irreversible parts of execution conditional on checks that can still be explained to the operator.

4.3 The run-set allocation

run_batch in 5. Experiment/1. Harness/scripts/run_batch.py lines 913 to 957 validates duplicate cell identities, rejects more than one model, applies the specification environment, rejects conflicting resume and destination arguments, and either checks the existing cells/ directory or calls _create_run_set lines 156 to 171. The allocator holds publish_lock while it finds the greatest existing numeric prefix, chooses the next number, rejects an existing explicit destination, and copies the first 001-* template. It then calls _write_active_batch lines 543 to 555 with status running.

The design reason in the comments is isolation and recoverability. A cell folder name does not include the model or prompt variant, so duplicates could overwrite evidence. A run-set number is never reused. The active marker is written before the first cell so the interface has a live status even while the first subprocess is starting.

4.4 The chain preparation

The batch loads chain declarations through load_chains in 5. Experiment/1. Harness/scripts/task_sequences.py lines 127 to 192. The loader reads task manifests, builds predecessor and successor maps, and refuses circular or branching declarations. For each cell, run_batch calls _chain_start_tree in 5. Experiment/1. Harness/scripts/run_batch.py lines 230 to 270 and later calls _note_chain_outcome lines 280 to 310.

A chain is a consecutive sequence of tasks in which a later task starts from the tree produced by the earlier task. validate_sequence_order in 5. Experiment/1. Harness/scripts/task_sequences.py lines 222 to 289 is normally applied by preflight, which refuses a missing predecessor, wrong order, or unrelated cell between steps. During execution, _chain_start_tree supplies the predecessor workspace only when the predecessor reached an acceptable state. If it did not, _log_not_run in 5. Experiment/1. Harness/scripts/run_batch.py lines 227 to 235 records the refusal in execution_log.md and the current cell becomes a skipped result. This prevents a later task from being described as a small extension when it was actually asked to rebuild the missing feature.

4.5 The first cell invocation

For a cell that is not already terminal, _run_one_cell in 5. Experiment/1. Harness/scripts/run_batch.py lines 449 to 540 creates the cell directory and a batch-phase timer. Unless resume logic says otherwise, its first stage is driver_run_cell at lines 490 to 495, which invokes run_cell.py with _run_cell_args lines 179 to 215. The next stage is driver_score_cell at lines 500 to 502. The driver reads metrics.json through _invocation_errored lines 464 to 468 before scoring, because an agent invocation error is an invalid run rather than a scoreable product. It also reads the reached marker through _reached_without_an_agent lines 475 to 483 and returns a reached record without scoring.

The two child programs are deliberately behind invoke. The comment at default_invoke explains that cell output streams directly while closeout output is captured for progress evidence. The driver therefore measures the outer duration of both child programs while leaving their own records under the cell directory.

4.6 The hidden-failure resume

After the initial score, _run_one_cell lines 504 to 540 reads scoring_summary.json and metrics.json. It returns when the hidden check passes, the visible check fails, the scorer is missing, the budget is exhausted, the maximum passes have run, or the resume count has reached its guard. Otherwise it invokes run_cell.py again with --resume-hidden-fail at lines 529 to 533, scores again at lines 536 to 539, increments the resume count, and evaluates the same conditions.

The reason stated by the code is to repair a hidden failure only when the visible result is green and the experiment still has a pass and budget headroom. The retry notice remains owned by run_cell.py; the batch driver supplies only the flag. A resumed invocation error returns the error record immediately, so a failed retry cannot be mistaken for a scored outcome.

4.7 The batch loop and circuit breaker

run_batch lines 979 to 1044 walks the specification’s cells in order. It skips an already terminal cell during a resume, reconstructs chain state while doing so, handles a refused chain step by logging it, and appends each returned cell dictionary. After a cell, _note_chain_outcome records whether its chain can continue. Two consecutive records with agent_invocation_error set mark aborted and append batch_aborted_api_errors for every remaining cell at lines 1034 to 1041. A non-error cell resets the counter.

The circuit breaker is not an ordinary scoring failure rule. The comment identifies repeated invocation errors as evidence that the environment, credit, or authentication is broken, so spending the remaining cells would produce no useful verdicts. If the loop reaches its end, finished is true and the eventual batch state is complete unless the breaker set aborted. If a base exception escapes, the finally block classifies it as stopped before completion and still performs closeout.

4.8 The closeout chain

_closeout in 5. Experiment/1. Harness/scripts/run_batch.py lines 833 to 850 takes the campaign publication lock and calls _closeout_locked lines 852 to 895. The locked function first writes active_batch.md, batch_outcome.json, and the run-index status, then creates closeout_progress.json. It iterates the declared chain of aggregate_metrics.py, update_monitoring.py, gen_iteration_summary.py, gen_run_set_report.py, gen_run_set_analysis.py, cleanup_run_set.py, and publish_experiment_results.py at lines 860 to 892. Each stage is marked running before invocation and done, skipped, or failed afterward. A child failure sets the overall progress status to failed but does not prevent the next stage from being attempted.

The reason is explicit in the comments: status records are cheap and must exist before child programs run, the progress file is polled by readers, and a failure in one derived artifact must not erase the evidence already produced by other stages. _write_closeout_progress writes a temporary file, flushes and synchronizes it, then replaces the visible JSON path at lines 685 to 718. This gives readers either the prior complete record or the new complete record.

4.9 The concurrent lane path

load_lane_specs in 5. Experiment/1. Harness/scripts/lane_coordinator.py lines 448 to 493 reads a non-empty lanes array and validates the required owner, endpoint, run identifier, batch, run set, and model fields. run_lanes lines 259 to 387 rejects duplicate run identifiers and duplicate destinations before spawning anything. For each lane it calls build_lane_env lines 107 to 133, records a start timestamp, and calls default_lane_invoke lines 199 to 223. The latter validates the batch endpoint and model, rejects forbidden argument overrides, and starts run_batch.py with a lane-specific environment supplied through Popen.

The coordinator then polls every started child in run_lanes lines 325 to 339. It does not wait for one lane to finish before inspecting another. When all have ended, it constructs the lane array with return code, timing, endpoint, model, run set, launch wait, and error fields at lines 340 to 371. For explicit run-set destinations it sums lease_wait_seconds from cell metrics and writes lane_provenance.json through write_lane_provenance lines 406 to 422. main lines 496 to 519 optionally writes a coordinator-level provenance file, prints the JSON record, and returns zero only when every lane row is successful.

4.10 The state machine rendering

The first diagram is the chapter 02 subgraph of 5. Experiment/11. Detailed Design/flow-model/flow_model.v001.json, rendered as Mermaid and limited to the batch and cell-driver states named by the model. The model identifiers used here are driver.run_cell, driver.invocation_check, driver.reached_check, driver.score_cell, driver.resume_decision, driver.resume_run_cell, driver.resume_invocation_check, driver.resume_score_cell, and driver.cell_result. The batch and closeout states are hand-drawn from the cited run_batch.py finalisation lines because the plan row explicitly names those records and the model subgraph does not enumerate them as separate nodes.

stateDiagram-v2
    [*] --> running: run_batch.py 957
    running --> complete: cell loop finished, run_batch.py 1045 to 1058
    running --> aborted_api_errors: two invocation errors, run_batch.py 1034 to 1041
    running --> stopped_before_completion: exception or stop, run_batch.py 1046 to 1063
    complete --> closeout_running: finally, run_batch.py 1049 to 1064
    aborted_api_errors --> closeout_running: finally, run_batch.py 1049 to 1064
    stopped_before_completion --> closeout_running: finally, run_batch.py 1049 to 1064
    closeout_running --> complete: all stages done, run_batch.py 893 to 895
    closeout_running --> failed: one or more stages failed, run_batch.py 867 to 895
    complete --> [*]
    failed --> [*]

The state values running, complete, aborted_api_errors, and stopped_before_completion are the batch dispositions written by _write_batch_outcome at lines 625 to 646. The closeout values running, complete, and failed are held in the progress record created at lines 658 to 682 and finalised at lines 893 to 895. The diagram shows the closeout complete name in two roles: as a batch disposition before closeout and as the successful closeout progress status. The records distinguish them by file.

The model’s cell subgraph is rendered separately here so its node identifiers remain visible:

stateDiagram-v2
    [*] --> driver_run_cell: driver.run_cell, run_batch.py 491 to 495
    driver_run_cell --> driver_invocation_check: run_batch.py 496 to 497
    driver_invocation_check --> driver_cell_result: invocation error, 464 to 473
    driver_invocation_check --> driver_reached_check: no invocation error, 498 to 499
    driver_reached_check --> driver_cell_result: reached, 475 to 488
    driver_reached_check --> driver_score_cell: ordinary cell, 500 to 502
    driver_score_cell --> driver_resume_decision: score record, 504 to 520
    driver_resume_decision --> driver_cell_result: terminal, 520 to 527
    driver_resume_decision --> driver_resume_run_cell: hidden failure with budget, 529 to 533
    driver_resume_run_cell --> driver_resume_invocation_check: 534 to 535
    driver_resume_invocation_check --> driver_cell_result: invocation error, 534 to 535
    driver_resume_invocation_check --> driver_resume_score_cell: no invocation error, 536 to 539
    driver_resume_score_cell --> driver_resume_decision: resume count increment, 540
    driver_cell_result --> [*]

This second rendering names the model identifiers driver.run_cell, driver.invocation_check, driver.reached_check, driver.score_cell, driver.resume_decision, driver.resume_run_cell, driver.resume_invocation_check, driver.resume_score_cell, and driver.cell_result in Mermaid state labels converted to readable state names. The branch descriptions follow the model descriptions and the cited implementation lines.

5. The state machine

The batch state machine has two nested records. The batch outcome is the durable account of whether the cell loop completed and why it stopped. The closeout progress record is the durable account of whether the derived publication stages completed. A batch can be aborted_api_errors while closeout is complete, because the circuit breaker stops new cells but does not prevent aggregation of the cells already recorded. A batch can be stopped_before_completion while closeout is failed, because a later closeout child can fail after the stop has been recorded.

The transitions are guarded by the line ranges below. A normal completed cell loop sets finished at run_batch.py line 1045. The circuit breaker sets aborted at lines 1034 to 1041. An exception is captured and re-raised at lines 1046 to 1048, and the finally block chooses the stopped state at lines 1049 to 1064. Closeout stage failure is accumulated at lines 867 to 892 and becomes the progress status at lines 893 to 895.

stateDiagram-v2
    state "batch running" as B_RUNNING
    state "batch complete" as B_COMPLETE
    state "batch aborted_api_errors" as B_ABORTED
    state "batch stopped_before_completion" as B_STOPPED
    state "closeout running" as C_RUNNING
    state "closeout complete" as C_COMPLETE
    state "closeout failed" as C_FAILED
    [*] --> B_RUNNING: create or resume run set, run_batch.py 913 to 957
    B_RUNNING --> B_COMPLETE: finished and no abort, 1045 to 1060
    B_RUNNING --> B_ABORTED: consecutive_errors >= 2, 1034 to 1041
    B_RUNNING --> B_STOPPED: BaseException, 1046 to 1063
    B_COMPLETE --> C_RUNNING: finally calls _closeout, 1049 to 1064
    B_ABORTED --> C_RUNNING: finally calls _closeout, 1049 to 1064
    B_STOPPED --> C_RUNNING: finally calls _closeout, 1049 to 1064
    C_RUNNING --> C_COMPLETE: failed is false, 893 to 895
    C_RUNNING --> C_FAILED: failed is true, 867 to 895
    C_COMPLETE --> [*]
    C_FAILED --> [*]

This diagram is hand-authored from the state fields and transitions in run_batch.py; it is not an additional flow-model claim. The model-rendered cell state diagram appears in section 4.10 and names the model nodes.

6. The sequence of one unit of work

The sequence below follows one cell in one lane. Programs are participants, the batch and cell files are records, and the Fabric service is an external service reached by the child runner rather than directly by the driver. The thread participant is the child process boundary in the batch driver’s invocation. The diagram is hand-authored from run_batch.py lines 490 to 539, lane_coordinator.py lines 306 to 339, and the cited preflight boundary. It is not a claim that the coordinator starts one thread per lane.

sequenceDiagram
    participant Operator as launch_batch.sh or launch_lanes.sh
    participant Coordinator as lane_coordinator.py
    participant Driver as run_batch.py child
    participant Cell as run_cell.py
    participant Scorer as score_cell.py
    participant Fabric as Fabric service
    participant CellFiles as cell metrics and scoring files
    participant Closeout as closeout programs

    Operator->>Coordinator: supply explicit lane plan
    Coordinator->>Driver: Popen with copied lane environment
    Driver->>CellFiles: create cell directory and stage timer
    Driver->>Cell: invoke driver_run_cell
    Cell->>Fabric: preflight or lease and inference through backend
    Fabric-->>Cell: admission, lease, and outcome
    Cell->>CellFiles: write metrics and execution records
    Driver->>CellFiles: read metrics for invocation and reached checks
    Driver->>Scorer: invoke driver_score_cell
    Scorer->>CellFiles: write scoring_summary.json
    Driver->>CellFiles: read score and budget state
    alt hidden failure with visible pass and budget headroom
        Driver->>Cell: invoke --resume-hidden-fail
        Cell->>Fabric: resume pass and inference
        Fabric-->>Cell: outcome
        Cell->>CellFiles: update metrics
        Driver->>Scorer: invoke resumed scorer
        Scorer->>CellFiles: update scoring_summary.json
    else terminal or refused cell
        Driver->>CellFiles: retain error, reached, or skip evidence
    end
    Driver->>Closeout: invoke stages under publish_lock
    Closeout->>CellFiles: read cell records
    Closeout->>CellFiles: write campaign and run-set outputs

The coordinator’s polling is concurrent at the process level. Within each child, cell order remains sequential because run_batch.py owns the cell loop. The external Fabric calls are owned by the cell runner and its backend, so this chapter records the boundary and its effect on the driver without attributing lease internals to the batch coordinator.

7. The records

The canonical rows below are records written by this mechanism or records that the mechanism directly orchestrates. A record is a durable file or table row, not an in-memory return value. The cell runner and scorer own the canonical details of metrics.json and scoring_summary.json; they are included here because the batch driver reads them to make control decisions. The campaign aggregator and interface read the resulting files after closeout.

RecordWriter and lineFields or contentReaders
active_batch.mdrun_batch.py _write_active_batch, lines 543 to 555Run set name, cell count, status, and timestamp wordingResults interface and monitoring readers, spread across their status views
run_index.csv status rowrun_batch.py set_run_index_status, lines 773 to 831The matching run_id row’s status; all other columns and rows are preservedCampaign run-index monitor and interface readers
batch_outcome.jsonrun_batch.py _write_batch_outcome, lines 617 to 650schema_version, record_provenance, run_id, status, stopped_before_completion, reason, cells_planned, cells_attempted, cells_skipped, cells_not_reached, and recorded_atattest_batch_outcome.py build_record lines 91 to 138 when repairing, report and campaign readers, and the interface’s run-set views
closeout_progress.jsonrun_batch.py _write_closeout_progress, lines 685 to 718, called by _closeout_locked lines 857 to 895Schema version, run id, batch status, overall status, timestamps, and per-stage key, title, script, status, times, seconds, message, and artifacts_closeout_locked on its own next write, monitoring readers, and the interface’s closeout view
execution_log.mdrun_batch.py _log_not_run called at lines 1015 to 1023Timestamped or formatted refusal reason for a chain cell not runRun-set reports and readers explaining skipped cells
cells/<cell_id>/metrics.jsonrun_cell.py, canonical writer outside this chapter; read by _invocation_errored and _reached_without_an_agent in run_batch.py lines 464 to 483 and by the resume decision lines 514 to 516Invocation-error marker, reached status, passes run, budget state, and cell execution measurementsrun_batch.py, score_cell.py, campaign aggregation, and interface readers
cells/<cell_id>/scoring_summary.jsonscore_cell.py, canonical writer outside this chapter; read by _run_one_cell lines 506 to 527Hidden pass, visible pass, missing-scorer marker, score states, and invalid reasonsrun_batch.py, campaign aggregation, and interface readers
lane_provenance.jsonlane_coordinator.py write_lane_provenance, lines 406 to 422, called from run_lanes lines 372 to 386Coordinator record plus the lane row: owner, endpoint, run id, model, batch, run set, return code, success, timestamps, launch wait, Fabric lease wait when metrics exist, contention, and errorCampaign run-index monitoring and interface readers
Optional coordinator provenance filelane_coordinator.py main, lines 511 to 518The rendered coordinator JSON record for every laneOperator tools and monitoring readers named by the invocation
Shared lane sinklane_coordinator.py publish_lane_results, lines 425 to 445Existing JSON rows plus copied lane result rowsThe coordinator-level monitor or aggregator that owns the supplied sink path
Lock sentinelpublish_lock.py publish_lock, lines 31 to 55No experiment data. The file is an operating-system lock targetConcurrent driver and coordinator publication processes

The lineage from a specification to campaign tables is shown below. The structural diagram is hand-authored as required by the diagram plan. It cites the specification fields at run_batch.py lines 179 to 216 and the run-set allocation and closeout regions at lines 913 to 1064.

flowchart LR
    Spec[batch specification: project, cells, model, fabric, budgets]
    Plan[lane plan: owner, endpoint, run id, batch, run set]
    Driver[run_batch.py child]
    Folder[run set folder]
    CellFolder[cells and execution_log.md]
    Metrics[metrics.json]
    Scores[scoring_summary.json]
    Outcome[batch_outcome.json]
    Progress[closeout_progress.json]
    Lane[lane_provenance.json]
    Tables[6. Metrics campaign tables]
    Monitor[7. Monitoring and interface views]
    Reports[reports and publication]
    Plan --> Driver
    Spec --> Driver
    Driver --> Folder
    Folder --> CellFolder
    CellFolder --> Metrics
    CellFolder --> Scores
    Driver --> Outcome
    Driver --> Progress
    Plan --> Lane
    Folder --> Lane
    Metrics --> Tables
    Scores --> Tables
    Outcome --> Reports
    Tables --> Monitor
    Progress --> Monitor
    Lane --> Monitor
    Tables --> Reports

The diagram separates canonical cell records from derived campaign tables. The closeout programs, not the batch driver itself, perform the transformation from cell and score files into the campaign tables. The driver owns the ordering and status boundary that makes that transformation possible.

8. The loops and the waits

The cell loop is the for c in cells loop in 5. Experiment/1. Harness/scripts/run_batch.py lines 980 to 1044. Its bound is the number of cells in the loaded specification. It stops normally after the last cell, or early when the consecutive invocation-error count reaches two. On resume, terminal cells are visited for state reconstruction and recorded as already_complete rather than staged again.

The per-cell scoring loop is the while True loop in _run_one_cell lines 504 to 540. Its normal bound is the remaining pass and token budget, represented by passes_run < max_passes and budget_exhausted in the metrics record. Its explicit safety bound is resumes < max_passes. It stops and returns when the hidden result is true, the visible result is false, the scorer is absent, the budget is exhausted, the pass limit is reached, or a child reports an invocation error. It retries only the hidden-failure case behind a visible pass with headroom.

The closeout loop is the for index loop in _closeout_locked lines 860 to 892. Its bound is the fixed CLOSEOUT_CHAIN tuple at lines 564 to 595. It stops after all seven declared stages have been attempted. A failed stage alters its own status to failed and sets the aggregate failed flag, but the loop continues.

The lane launch loop is the for spec in lane_specs loop in run_lanes lines 306 to 323. Its bound is the number of lane specifications. A launch exception records one failed LaneResult and continues to the next lane. The lane polling loop is while pending at lines 325 to 339. Its stop condition is an empty pending set. It calls proc.poll() for each pending process and sleeps for poll_seconds, default 0.05, only when at least one process remains.

The lock wait is unbounded and blocking. publish_lock in 5. Experiment/1. Harness/scripts/publish_lock.py lines 37 to 55 calls fcntl.flock with LOCK_EX; no timeout or retry count is configured. The purpose is to serialize read-modify-write publication on the same host. It releases the lock in finally, including when the protected closeout raises.

Preflight has no batch-level retry loop. The host probe performs one acquire, one heartbeat, one protocol probe, and one release in probe_dgx_host lines 249 to 276. The container probe delegates to one Docker subprocess with a timeout of 1200 seconds at lines 329 to 332. Controller admission retries, when any, belong to the Fabric lease backend and are outside the owning files of this chapter.

The closeout file write has no repeated attempt after an operating-system error. _write_closeout_progress reports the error, removes its temporary file when possible, and returns false at lines 710 to 718. The caller continues the stage loop. The batch outcome write similarly reports a disk error at lines 643 to 650 and returns the in-memory payload.

9. The guards and refusals

Guard locationCheckRefusal or alteration
run_batch.py lines 918 to 931Cell identities resolve to distinct folder namesRefuses duplicate cells before run-set creation.
run_batch.py lines 932 to 938At most one model is namedRefuses a mixed-model batch because the folder identity lacks the model.
run_batch.py lines 940 to 941Resume and explicit destination are not both suppliedRefuses conflicting run-set modes.
run_batch.py lines 949 to 953Resume path contains cells/Refuses an invalid resume target.
run_batch.py lines 160 to 170A 001-* template exists and destination does not existRefuses allocation without a template or on a destination collision.
preflight_batch.py lines 383 to 444Task manifests are readable and fit, and sequence declarations are validRefuses unfit or unknown tasks and invalid chain order.
preflight_batch.py lines 477 to 499Cell identities are unique and models are not mixedRefuses the static plan.
preflight_batch.py lines 190 to 204Fabric lane is present and authorizedRefuses a missing or unknown lane.
preflight_batch.py lines 213 to 235Fabric is not in maintenance, has lane capacity, and contains the aliasRefuses a host probe that cannot admit the planned work.
preflight_batch.py lines 241 to 244Backend and model capabilities are compatibleRefuses an incompatible model route.
preflight_batch.py lines 310 to 345Cell image and container bridge probe succeedRefuses a container path that cannot execute the measured cell.
run_batch.py lines 520 to 527Hidden failure, visible pass, budget, pass count, scorer presence, and resume countReturns the current result rather than spending another resume when any required condition is absent.
run_batch.py lines 1034 to 1041Two consecutive invocation errorsMarks the batch aborted_api_errors and marks remaining cells skipped.
lane_coordinator.py lines 292 to 299Lane run identifiers and destinations are uniqueRefuses the whole launch before any child starts.
lane_coordinator.py lines 234 to 256Batch endpoint and model agree with lane provenanceRefuses the individual launch.
lane_coordinator.py lines 211 to 219Extra arguments do not replace protected paths and mode is new or resumeRefuses an unsafe child command.
run_tranche.sh lines 73 to 89Lane identity is complete, or no scoped lane containers are active in legacy modeRefuses ambiguous or broad cleanup.
publish_lock.py lines 49 to 55Shared publication enters the exclusive lockBlocks rather than interleaving updates.
run_batch.py lines 93 to 99SIGTERM has not already been forwardedForwards the first stop to the child and ignores repeats.

These checks alter flow at different layers. A preflight refusal prevents a run set. A batch validation refusal prevents cell staging. A chain refusal leaves a skip record but permits later independent cells. A circuit-breaker decision leaves an outcome and skips the unprocessed suffix. A closeout child refusal changes a stage to failed but allows later stages to run.

10. Every unhappy path

The four parts in each row are the intended step, the reason for its design, the trigger and record, and the cost and downstream effect. The driver column names the batch driver’s action, the scorer column says whether scoring is reached, the aggregator column describes available evidence, and the interface column describes the visible status.

Trigger and code pathIntended step and design reasonRecord and status or exit codeCost and downstream effect
Preflight failure in run_batch.py lines 1098 to 1105Validate spendability, cost, task declarations, or endpoint before staging because a refused plan should not create a half-run set.JSON refusal on standard output, rendered details on standard error, exit code 1, and no run-set record.The driver does not start, the scorer and aggregator receive no cell records, and the interface can report the preflight refusal rather than a running batch.
Duplicate cell identity or mixed model in run_batch.py lines 918 to 938Allocate non-overlapping cell folders because overwriting a folder destroys factor provenance.SystemExit before _create_run_set; no batch outcome and no run set from this invocation.Only parsing and validation time is spent. The scorer and aggregator have nothing to consume, and the interface sees a launch error rather than a batch.
Invalid resume path or conflicting destination in run_batch.py lines 940 to 953Re-enter only a recognizable run set and choose one allocation mode because a new and resumed run cannot share semantics.SystemExit; no new run-set record.No cell is launched. Downstream programs see no new evidence; the interface receives the command failure.
Missing template or destination collision in _create_run_set, lines 156 to 171Create a unique isolated folder under the allocation lock because the cell loop requires a stable destination.SystemExit; no copied run set.Disk and lock time are spent. No scorer, aggregator, or interface batch record is produced.
Unfit, unreadable, or unknown task in probe_task_registrations, lines 348 to 445Refuse tasks whose verdict cannot be interpreted as a measured result because the task declaration is part of the design.Preflight ok false, exit code 1 through main; no run set.The cost is manifest inspection. No driver cell loop or scorer is reached, so the aggregator and interface have only the refusal report.
Broken chain declaration or order in load_chains and validate_sequence_order, task_sequences.py lines 127 to 289Carry a chain from its head in contiguous order because a later step otherwise measures reconstruction rather than the registered addition.Preflight refusal with violation text, exit code 1; no run set.File reads only. No driver, scorer, aggregator, or run-set interface record is created.
Fabric maintenance, zero capacity, absent alias, incompatible backend, or host probe failure in preflight_batch.py lines 213 to 289Confirm the selected machine can serve the selected protocol before the first measured cell because queueing a known-impossible batch spends time without evidence.Preflight JSON or text refusal, exit code 1; the container probe is skipped when the host probe fails.One host probe may have contacted the controller. The driver and scorer do not run, and the interface sees a preflight refusal rather than a live run.
Container probe failure in preflight_batch.py lines 292 to 345Test the actual measured-cell network route because a host-only success does not prove container connectivity.Preflight refusal, exit code 1, with the subprocess detail; no run set.Docker probe time is spent. No cell, score, aggregation, or normal interface result exists.
Chain predecessor did not pass in _chain_start_tree, run_batch.py lines 1013 to 1023Prevent a dependent task from starting on a tree missing its prerequisite.execution_log.md refusal and a cell result with skipped; batch processing continues to later cells.Only the dependency check and log write are spent. The scorer is not invoked for the skipped cell, the aggregator can count the skip, and the interface can show the reason.
Agent invocation error in _run_one_cell, run_batch.py lines 464 to 497 or 534 to 535Avoid scoring a process that did not produce a valid agent run because an API or launch error is not a product verdict.Returned cell record has agent_invocation_error: true, visible_pass: false, and no scoring invocation for that attempt.Time and any partial cell cost are lost. The driver increments its consecutive-error count, the scorer is bypassed, the aggregator receives an invalid cell record, and the interface can show an invocation failure.
Two consecutive invocation errors in run_batch.py lines 1034 to 1041Stop an environment-wide failure before it consumes the remaining plan because repeated API errors are not independent scored failures.batch_outcome.json status aborted_api_errors; remaining cells have skipped: batch_aborted_api_errors; normal process exit remains the completed driver path.Completed cells still reach closeout and aggregation. The scorer never sees the skipped suffix, the aggregator can distinguish attempted and skipped cells, and the interface shows an aborted batch.
Hidden failure without the required visible pass, budget, pass, or scorer conditions in _run_one_cell, lines 518 to 527Resume only a recoverable hidden failure because an absent scorer or failed visible suite cannot be repaired by blind repetition.The current cell result is returned with hidden, visible, missing-scorer, resume, and budget fields; no extra invocation.No retry cost is incurred. The scorer has its existing summary, the aggregator retains the advisory or invalid fields, and the interface shows the terminal cell state.
Resume invocation error in _run_one_cell, lines 529 to 535Treat a failed recovery attempt as an invocation error rather than a new scored verdict.Error record with the number of attempted resumes; the batch may then trip its circuit breaker.Retry time and partial usage are spent. The scorer is not called for that retry, the aggregator sees the error, and the interface can distinguish it from a hidden failure.
Stop signal in _forward_sigterm, run_batch.py lines 79 to 99Stop the child once and enter the driver’s finally path because an orphaned child can continue spending or hold resources.Child receives termination, the driver raises SystemExit(143), and closeout writes stopped_before_completion with a reason.Work in the active child may be lost, but completed cells are still aggregated. The scorer only has records already written, the aggregator runs during closeout, and the interface leaves a stopped status instead of running.
Unhandled exception in the cell loop, run_batch.py lines 1046 to 1064Preserve the exception while still finalising the run because a crash must not erase completed cells.Exception is re-raised after closeout; batch_outcome.json records stopped_before_completion and the derived reason.The active cell may lack a complete record. The scorer and aggregator consume only durable records, and the interface sees a stopped run with an explanation.
Closeout child returns nonzero or raises in _closeout_locked, lines 867 to 892Attempt every derived stage because one missing report should not suppress monitoring or publication of other valid artifacts.The stage is failed, its message is recorded, and final closeout_progress.json is failed; the driver continues the chain.The failed artifact may be absent, but later stages get their chance. The aggregator may be complete or failed depending on the stage, the scorer is unaffected, and the interface can show the exact failed stage.
Closeout progress or batch outcome disk write fails, lines 643 to 650 and 710 to 718Make status writes best effort without replacing the original closeout failure or preventing child stages.Error text is printed to standard error; the in-memory closeout continues, with the affected file absent or stale.The evidence boundary is weakened. Existing cell records remain available to the scorer and aggregator, but the interface may show stale or missing progress.
Duplicate lane run id or destination in run_lanes, lane_coordinator.py lines 292 to 299Prevent two children from targeting one identity or folder because concurrent resume work would collide.ValueError before any lane launch; no coordinator result record.Only validation time is spent. No driver or scorer runs, and the interface receives a coordinator launch failure.
Invalid lane plan or unknown endpoint in load_lane_specs, lines 458 to 487Require a complete explicit plan because automatic discovery would make the concurrent experiment non-reproducible.ValueError; no child is started and no lane provenance is written.Plan parsing only. The aggregator has no lane row and the interface sees a plan error.
Lane batch endpoint or model mismatch in _validate_lane_batch_endpoint, lines 234 to 256Keep labels and provenance truthful because a lane record must identify the machine and model actually selected.ValueError from the lane launch path; that lane has no child record, while validation failure in main aborts the coordinator command.No cell cost is incurred. No scorer or aggregator record exists for the refused lane, and the interface sees the consistency refusal.
Lane child launch exception in run_lanes, lines 306 to 318Isolate one launch failure so a valid lane can still run.A LaneResult has null return code, end time, wait time, and an error string; other lanes are polled.The failed lane spends launch time only. Its driver and scorer do not run, while other lanes aggregate normally; the coordinator and interface can show one failed lane beside another result.
Lane child exits nonzero in run_lanes, lines 325 to 369Wait for every child independently because one lane’s failure must not terminate another lane.The lane row has its nonzero return code, ok: false, end time, and any error; the coordinator returns exit code 1 if any row is not ok.The failed lane’s own closeout may still have produced records. The other driver and scorer remain independent, aggregation is per run set, and the interface shows the lane result.
Lane cancellation in cancel_lane, lane_coordinator.py lines 390 to 403Terminate only the selected lane and its containers because broad cleanup could kill another endpoint’s cell.The selected process is terminated and its scoped containers are stopped; other lane records are unchanged.The cancelled lane may be incomplete and later closeout depends on its child behavior. Other scorers and aggregators continue, while provenance preserves the lane distinction.
Lock contention in publish_lock.py lines 49 to 55Serialize shared publication because read-modify-write without a lock loses one lane’s update.No failure record is created; the caller blocks until the lock is available and then enters its critical section.Wall time increases by the other publisher’s hold time. Driver closeout, aggregation, and interface updates wait but do not interleave.

11. The metrics produced or fed

The batch driver produces control metrics and feeds measurement metrics. The outer cell timer created at run_batch.py lines 461 to 462 records the duration of the driver-run and driver-score stages. The cell runner’s metrics.json remains the canonical source for passes, budget exhaustion, invocation errors, and measured cell values. The scorer’s scoring_summary.json supplies hidden and visible verdicts, score states, and invalid reasons. The batch driver’s returned cell dictionary carries the control projection used for chain bookkeeping and the circuit breaker.

Source fieldProgram and routeDestination or useStatus
passes_run, budget_exhausted, and invocation markersrun_cell.py writes metrics.json; run_batch.py reads lines 464 to 516Cell row inputs to the campaign aggregator and resume guardMeasurement or validity field, not merely advisory
hidden_pass, visible_pass, and hidden_scorer_missingscore_cell.py writes scoring_summary.json; run_batch.py reads lines 506 to 527Cell verdict columns and hidden-failure resume decisionVerdict fields; missing scorer is an instrument failure
resumes_run_one_cell lines 504 to 540Returned cell result and batch reportAdvisory count of recovery attempts
agent_invocation_error_error_record lines 470 to 473Circuit-breaker count and invalid cell recordValidity field that alters batch flow
cells_planned, cells_attempted, cells_skipped, and cells_not_reached_write_batch_outcome lines 625 to 640Report and campaign run dispositionControl and coverage metrics
Stage status, seconds, message, and artifacts_closeout_locked lines 860 to 895closeout_progress.json, monitoring, and interfaceOperational metrics, advisory to scientific scoring
Lane wait_seconds and makespan_secondsrun_lanes lines 340 to 369lane_provenance.json and throughput analysisAdvisory scheduling metrics
Summed lease_wait_seconds from cell metricsrun_lanes lines 377 to 386Lane contention field in provenanceAdvisory contention metric; lease accounting is owned by the cell lifecycle
Run-index statusset_run_index_status lines 806 to 829Campaign run index and interface statusOperational disposition, not a score

The closeout stage aggregate_metrics.py turns the cell records into campaign tables; this chapter does not invent the table schema or claim ownership of those columns. It records which source fields the batch driver guarantees are present or uses to decide whether the aggregation can proceed. Cost estimates from preflight are advisory planning values unless the preflight result is retained by the caller; they do not become a measured cell metric.

12. The tests

The verified test inventory exercises the mechanism through seams and temporary run-set fixtures. 5. Experiment/1. Harness/scripts/tests/test_run_batch.py covers ordinary cell iteration, hidden-failure resume, invocation errors, run-set allocation, and the injected invocation boundary. test_run_batch_sequences.py covers chain order, refusal, and continuation. test_chained_gate_satisfied.py and test_reached_step_scoring.py cover a step satisfied by its start tree and the rule that a reached cell is not scored. test_resume_into.py covers terminal-cell skipping, halted cells, and in-place resume. test_prompt_variant_and_retry_feedback.py covers forwarded specification factors and the duplicate and mixed-model refusals.

test_preflight_batch.py covers spendability, cost, host and container probe outcomes, lanes, maintenance, and static checks. test_task_registration_gate.py covers fit, unfit, and unreadable task manifests. test_task_sequences.py covers sequence construction, missing steps, wrong order, and scattered chain cells. test_fabric_endpoint_wiring.py covers endpoint selection and forwarding into both probes.

test_lane_coordinator.py covers scoped labels, copied environments, duplicate identities, endpoint and model consistency, overlapping execution, independent lane failure, cancellation, restart isolation, provenance, and concurrent publication. test_publish_lock.py covers the local lock. test_closeout_progress.py covers stage progress, failed stages, stopped batches, atomic-write failure, and continuation. test_run_set_report.py covers completed and stopped outcomes, run-index closure, attestation counts and discrepancies, and the distinction between driver-written and attested records. test_integration_smoke.py covers the connected record contract across the batch and aggregation seams.

The mapping from every unhappy path in section 10 to an individual test was not made. In particular, this chapter does not claim a named test for every disk-level error, every shell launch failure, every external Fabric response, or every interface rendering consequence. The available tests establish the principal control branches through fakes, but the path-by-path coverage mapping remains unchecked.

13. The dated incidents that shaped the code

The comments in the owning files preserve operational causes for guards that might otherwise look arbitrary. The dates and line locations below are included so a reader can distinguish a deliberate refusal from an accidental limitation.

On 2026-08-09, the run set 008 experience shaped the invocation-error guard. run_batch.py lines 465 to 468 state that a pass which died on an API error must not be scored, and lines 959 to 962 state that two consecutive such failures indicate a broken environment. The result is the agent_invocation_error record followed by the batch circuit breaker at lines 1034 to 1041.

On 2026-08-13, after run set 008, the preflight comments in run_batch.py lines 1089 to 1097 explain why the real command-line entry point owns the preflight call. It prevents creation of a run set that the environment cannot finish, while keeping the injectable run_batch function free of a live CLI probe. The corresponding spendability rationale is in preflight_batch.py lines 6 to 23 and 107 to 154.

On 2026-08-16, the preflight comments in preflight_batch.py lines 113 to 121 changed the spendability probe to use the same container wrapper as the measured cell whenever FIVEB_CELL_CONTAINER is set. The stated incident was a green host-side gate followed by a container authentication failure in run set 010. The container probe at lines 292 to 345 exists to test that actual route.

On 2026-08-29, an operator directive shaped in-place resume. run_batch.py lines 943 to 948 state that a stopped run set is re-entered rather than replaced, with terminal cells skipped and a halted cell restaged. The same date appears in the signal-handler comment at lines 79 to 90, which records the orphaned lease caused by a bare stop in run set 038 and explains why the first SIGTERM is forwarded to the active child while repeats are ignored.

On 2026-08-29, the specification endpoint rule was also fixed. run_batch.py lines 423 to 438 state that the batch specification is the source of Fabric levers and that old specifications without an endpoint continue to resolve to dgx. The preflight endpoint is resolved once and passed to both probes at preflight_batch.py lines 532 to 544 so host and container cannot silently test different machines.

On 2026-08-30, the chain comments in preflight_batch.py lines 385 to 403 describe twenty-one cells that ran from a frozen build without their predecessor feature. The code consequently validates the complete chain order and lets run_batch.py pass the predecessor tree through _chain_start_tree.

On 2026-08-31, closeout was moved into the finally path. The module comment at run_batch.py lines 17 to 24 names run sets 041 and the loss of completed verdicts, metric rows, and reports after an intentional stop. The implementation at lines 1049 to 1064 now writes the disposition and runs the closeout chain for both normal and abnormal exits. The same date appears in the shell launcher comments at launch_batch.sh lines 31 to 54, which moved specification commit work out of the web request, and in launch_lanes.sh lines 33 to 48, which serializes commits for concurrent launchers.

On 2026-09-01, the duplicate identity comments at run_batch.py lines 921 to 931 record a plan whose prompt variants resolved to the same eight folders for twelve intended cells. This is why prompt variant and retry-feedback differences cannot bypass the folder identity guard. The mixed-model explanation at lines 898 to 904 gives the same protection for model differences.

On 2026-09-03, the reached-step comment at run_batch.py lines 476 to 480 records tracker 2.35 and explains why a chained step satisfied by its start tree is returned without invoking the scorer. The corresponding chain state is defined in task_sequences.py lines 58 to 79.

The lane coordinator comments identify LP-16 as the concurrency issue addressed by the file. Its module docstring at lane_coordinator.py lines 2 to 37 states that two lanes must not share the interpreter whose environment mutation selects Fabric settings. The child environment and Popen boundary at lines 107 to 133 and 199 to 223 preserve that isolation. The lock module’s docstring at publish_lock.py lines 1 to 22 records the related read-modify-write race in run index, metric tables, and monitoring state. These comments explain why run_lanes can be concurrent while closeout publication remains serial under a local lock.

14. The weakest claim, what was not checked, and the token line

The weakest claim is the complete downstream effect of every closeout failure on the interface and campaign tables. The owning code proves that stages are attempted in order and that progress records distinguish done, skipped, and failed, but this chapter did not inspect every reader of every derived table or every external Fabric response. The one-to-one mapping from each unhappy path to a test was not made, and the exact closeout table schemas remain owned by their aggregators. The diagrams also include hand-authored batch and lineage structure around the explicitly identified flow-model cell nodes. These are the limits of verification for this chapter.

Model: openai/gpt-5.6-luna via OpenRouter; tokens: see the job ledger