The flow model does not contain a chapter 07 subgraph in v001. The model records the ordinary cell and driver nodes, including cell.satisfied_by_start_tree, driver.reached_check, and driver.resume_decision, but its chapter entries stop at 06 and continue at 08. The state, sequence, and lineage diagrams in this chapter therefore use the verified source code and the plan row, not an absent model subgraph. The structural diagrams are hand-authored with cited elements. The local draft was treated as a lead only. The owning scripts and the tests named here were read at the verified commit.
1. The purpose and the position in the life of the experiment
A chained sequence is a set of maintenance tasks whose workspaces are intended to be carried forward in order. A workspace is the directory containing the product tree that one cell hands to the next cell. The first task, called the head, begins from the frozen build. A later task begins from the workspace left by its predecessor, so its result measures an addition to the already built feature rather than an attempt to recreate the whole sequence.
A reached step is a chained task whose predecessor workspace already satisfies the current task’s hidden checks. Hidden checks are the registered checks that are not shown to the coding agent. In this case the runner does not invoke an agent, spends no tokens, leaves the workspace unchanged, and records the step as satisfied_by_start_tree. That verdict is a successful chain state, not a skipped task. A resumed cell is an existing measured cell continued after a hidden failure while it still has passes and cumulative budget available. It keeps its workspace, staged snapshot, ledger, and pass numbering. A halted cell is different: it was interrupted part way through a run and its leftover workspace is not treated as safe to resume.
The batch driver invokes this mechanism while walking the cells in a batch. run_batch.py calls the sequence loader before the cell loop, chooses each chained cell’s start tree, refuses a later step after a predecessor failure, and decides whether an existing cell is terminal when resuming into a run set. The cell runner stages or continues the workspace, applies the staging gate, records a reached step when the gate says the predecessor tree already passes, or runs the agent and its pass loop. The scorer preserves a reached summary rather than rescoring it as an ordinary cell. The restoration utility repairs a historical case in which a scorer overwrote that summary. The mechanism produces a carried workspace, a chain verdict, cell metrics, a scoring summary, and an execution-log explanation that downstream aggregation and the interface can distinguish from an ordinary skip.
The purpose is therefore continuity with an honest cost record. Chain continuity prevents a later task from being measured against the wrong source tree. The reached record prevents an already fulfilled step from blocking the sequence or being charged as agent work. Resume rules preserve one measured cell across hidden-failure passes without turning an interrupted or exhausted cell into an apparently fresh attempt.
2. The reader’s map of the owning files
2.1 run_batch.py
run_batch.py is the batch orchestrator. Lines 1 to 226 contain imports, constants, JSON helpers, cell naming and process support. The chain key and source-tree region is run_batch.py lines 244 to 293, where _chain_key identifies a chain run by head, condition, and seed, and _chain_start_tree returns the frozen-build choice, the predecessor workspace, or a refusal. Chain outcome handling is run_batch.py lines 296 to 325, where _note_chain_outcome reads the scorer verdict and marks a chain broken unless the correctness state is pass or satisfied_by_start_tree. Halt and terminal predicates are run_batch.py lines 327 to 414: _cell_was_halted distinguishes an interrupted invalid run from an intentional validation run, and _cell_is_terminal decides whether resume mode may skip the cell.
The main cell flow is in the larger execution region beginning before the closeout at line 970 and continuing through the cell loop at run_batch.py lines 981 to 1040. It rebuilds chain state for completed cells, refuses a cell whose predecessor workspace is absent or whose chain is already broken, invokes the cell and scorer, then notes the outcome. The argument parser is at lines 1070 to 1084, including --batch, --skip-preflight, --max-usd, --resume-into, and --run-set-destination. Closeout and final batch status are later in the file, outside the chain predicates, and write the batch outcome and run-set status even when the cell loop stops.
2.2 run_cell.py
run_cell.py owns the per-cell staging, pass loop, resume decision, and final records. Its constants and general helpers occupy lines 1 to 945. The environment-controlled ruling is run_cell.py lines 946 to 958, chained_gate_ruling_enabled, which reads FIVEB_CHAINED_GATE_SATISFIED on every call and treats every value except exactly 0 as enabled. The runner’s argument parser is at lines 1508 to 1554.
The cell setup and result construction surround the resume decision. The resume predicate is run_cell.py lines 1829 to 1836: --resume-hidden-fail, an existing workspace, and at least one streamed transcript are all required. Carry-forward of earlier invalidity is lines 1838 to 1857. The refusal helper is lines 1859 to 1868. The reached-step writer is run_cell.py lines 1870 to 1940, _satisfied_by_start_tree; it writes the scoring summary, metrics, scoring output copy, reached log line, and final manifests without invoking an agent. The ordinary staging and pass flow follows. Reached metrics are defined at run_cell.py lines 3065 to 3119 by reached_step_metrics, and the reached execution-log writer is lines 3122 to 3132 in _log_reached. Finalisation helpers follow this region.
2.3 score_cell.py
score_cell.py contains the scorer and its command-line entry point. Its imports and general scoring helpers precede the chained baseline region. baseline_tree_for_cell is score_cell.py lines 409 to 469. It selects the head’s preserved staged snapshot for a later chain step, then the frozen tree recorded in the manifest, then the cell’s own staged snapshot. The reason is that a hidden comparison must see the whole chain addition, not treat the predecessor’s work as pre-existing on both sides of the comparison.
The reached-summary reader and its call site are score_cell.py lines 1233 to 1268. reached_step_summary reads scoring_summary.json and returns it only when score_states.correctness_gate is satisfied_by_start_tree; score_cell then keeps that summary unless --force-rescore-reached was supplied. Ordinary scoring and finalisation occupy the later regions of the file. Its parser is at lines 1597 to 1614 and accepts --run-set, --cell, --skip-evidence, --refresh-validity, --dry-run, --force-rescore-reached, and optional run-set positional arguments.
2.4 restore_reached_step.py
restore_reached_step.py is a small repair utility. Lines 1 to 33 contain the module contract, imports, and the reached-state constant. newest_reached_prior is lines 34 to 45 and searches preserved scoring_summary.prior-*.json files. _identity is lines 48 to 68 and reconstructs the cell identity from the folder and run manifest. restore is lines 71 to 107: it selects a prior reached summary, restores the current summary and scoring output when needed, creates missing reached metrics through run_cell.reached_step_metrics, and appends a reached execution-log line. The command-line parser and JSON report are lines 110 to 115.
2.5 task_sequences.py
task_sequences.py is the declaration and validation library. Lines 1 to 57 contain the module contract, paths, and constants. step_was_reached is lines 69 to 79 and permits only pass or satisfied_by_start_tree to carry a chain onward. Chains and its lookup methods occupy lines 92 to 125. load_chains is lines 127 to 192 and reads task manifests, ignoring unreadable manifests, while rejecting circular declarations and branches. Grouping is lines 195 to 219. validate_sequence_order is lines 222 to 289, order_cells is lines 292 to 324, and prune_orphan_steps is lines 327 to 364. These functions enforce head-first, consecutive order for one condition and seed, and can report or remove later steps with missing predecessors.
The owning files do not use numbered comment steps for a file over 50,000 bytes in this mechanism. The map above therefore uses their named functions and verified line regions rather than inventing a numbering scheme.
3. The inputs
3.1 Command-line arguments
The batch parser at run_batch.py lines 1073 to 1082 accepts --batch, the batch specification path; --skip-preflight, a flag; --max-usd, the optional cost ceiling; --resume-into, the existing run-set path; and --run-set-destination, the explicit destination for a new run set. The cell parser at run_cell.py lines 1510 to 1553 accepts the required run set, project, condition, task, and seed, plus phase, model, prompt variant, retry feedback, agent backend, maximum passes, token budget, dry-run, hidden-failure resume, base directory, Fabric endpoint, and start-tree options. The scorer parser at score_cell.py lines 1599 to 1613 accepts run set, cell, evidence and validity flags, dry-run, force-rescore-reached, and optional run-set paths. The repair parser at restore_reached_step.py lines 110 to 114 accepts a cell directory and --dry-run.
3.2 Environment variables
| Variable | Read at | Decision |
|---|---|---|
FIVEB_CHAINED_GATE_SATISFIED | run_cell.py lines 946 to 958 | Exactly 0 restores refusal of a chained step whose start tree passes; other values enable the reached ruling. |
FIVEB_FABRIC_ENDPOINT | run_cell.py lines 1721 to 1722 | Selects the Fabric endpoint when the command line does not name one. |
DGX_SPARK_FABRIC_MODEL | run_cell.py line 1682 | Supplies the model when --model is absent. |
DGX_SPARK_FABRIC_LANE | run_cell.py lines 1724 to 1725 and 2337 | Supplies the lane recorded for the cell and sent to the backend. |
DGX_SPARK_FABRIC_PRIORITY | run_cell.py line 2339 | Sets admission priority, defaulting to 50. |
DGX_SPARK_FABRIC_LEASE_SECONDS | run_cell.py lines 2340 to 2341 | Sets lease length. |
DGX_SPARK_FABRIC_REQUEST_PRESET | run_cell.py lines 2346 to 2347 | Selects the requested backend preset. |
DGX_SPARK_FABRIC_POLL_SECONDS | run_cell.py lines 2383 to 2384 | Sets admission polling interval. |
FIVEB_SKIP_EVIDENCE | score_cell.py line 1438 | Suppresses best-effort evidence capture after scoring. |
FIVEB_STAGE_NODE_MODULES | run_cell.py line 2010 | Alters staging of node modules. |
FIVEB_SKIP_STAGING_GATE | run_cell.py line 2167 | Alters the staging-gate path and therefore must be treated as a measurement control. |
Other environment values are passed to child processes or recorded in manifests, but the cited chain, reached-step, baseline, and resume decisions do not read them directly.
3.3 Files read
| File or source | Read at | Decision |
|---|---|---|
Task manifests under 4. Task Library/*/*/task_manifest.json | task_sequences.py lines 127 to 146 | Declares depends_on, from which the ordered chains are built. |
| Batch specification JSON | run_batch.py main flow and parser at lines 1073 to 1082 | Supplies cells, condition, seed, limits, and resume configuration. |
| Predecessor workspace | run_batch.py lines 284 to 293 | Supplies the start tree for a later chain step; absence breaks the chain. |
Prior cell_manifest.json and metrics.json | run_cell.py lines 1845 to 1857 and run_batch.py lines 375 to 413 | Carries invalidity, pass count, budget, and terminal decisions across resume. |
Prior scoring_summary.json | run_batch.py lines 310 to 324 and score_cell.py lines 1245 to 1254 | Supplies the correctness verdict that continues or breaks a chain. |
Preserved scoring_summary.prior-*.json | restore_reached_step.py lines 34 to 45 | Supplies a reached verdict for repair. |
| Run manifest | restore_reached_step.py lines 55 to 66 | Supplies project and model identity during repair. |
3.4 Network endpoints
The chain and reached decisions are local and do not call a network endpoint. If the cell is admitted to ordinary agent work, the backend used by run_cell.py may call the Fabric controller and the agent service, but those calls belong to the lease and cell-execution chapters. A reached step makes no agent request. The scorer and restoration utility are local programs; the scorer runs registered checks locally and the restoration utility writes local JSON and Markdown records.
4. The happy path in order
4.1 Sequence declaration loading
The batch loads chains with load_chains in task_sequences.py lines 127 to 192 before it walks cells. It scans task manifests for depends_on, builds predecessor and successor maps, walks backward to each head, and walks forward to construct each ordered chain. The comments give the design reason: declarations, rather than hard-coded task names, make adding or removing a dependency change the rule without changing the module. A circular declaration or a branch is refused because one produced tree cannot be carried into two successors.
The validation pass uses validate_sequence_order in task_sequences.py lines 222 to 289. It groups chain cells by head, condition, and seed, then reports a missing earlier step, wrong order, or unrelated cells between consecutive steps. The design reason in the comments is measurement meaning: a later task staged from the frozen build would be asked to recreate work it is registered only to extend. order_cells at lines 292 to 324 can rearrange the specification while preserving non-chain relative order. prune_orphan_steps at lines 327 to 364 can remove steps beyond the unbroken prefix and returns plain-language notes.
4.2 Start-tree selection
For each cell, _chain_start_tree in run_batch.py lines 257 to 293 determines the source directory. A standalone task and a chain head return no explicit tree, which means the frozen build. A later step constructs the predecessor cell identifier, checks that its workspace directory exists, and returns that directory. If the chain key is already broken, or if the predecessor workspace is absent, the function returns a refusal. The code comment gives the reason for refusing an absent workspace: silently falling back to the frozen build would reintroduce the defect the sequence machinery exists to prevent.
The driver applies this result in the cell loop at run_batch.py lines 1006 to 1023. A refusal is written through _log_not_run, added to the result list, and does not invoke either the runner or scorer. A valid start tree is passed to _run_one_cell at lines 1025 to 1027. The batch therefore has one explicit boundary between sequence control and cell execution.
4.3 Resume eligibility
When a batch resumes into an existing run set, _cell_is_terminal in run_batch.py lines 365 to 413 decides whether to skip an existing cell. It reads metrics and scoring records. A hidden pass, missing hidden scorer, visible failure, exhausted budget, or spent pass limit is terminal. A halted cell with an attempt remaining is deliberately not terminal, because its scorer may have judged a half-finished workspace and a fresh staging run is needed. The comments cite the design reason: judging that cell only by its verdict would permanently preserve an invalid result.
The driver rebuilds chain state while skipping terminal cells through _note_chain_outcome, called at run_batch.py lines 983 to 992. This is necessary because a resumed driver must know whether a skipped predecessor had failed before it allows a later cell to use a workspace. The skip_initial choice at lines 993 to 1004 carries an orderly existing scoring record into the cell runner but excludes halted cells.
4.4 Resume continuation
The runner evaluates the resume predicate at run_cell.py lines 1829 to 1836. The command flag, a workspace, and at least one pass*.stream.jsonl transcript are required. The comments call this continuation of the same measured cell: prior edits, staged snapshot, ledger, and pass numbering remain in place, and the token budget is cumulative. If prior metrics mark the cell invalid, lines 1838 to 1857 carry that invalidity forward rather than allowing a later finalisation to make the cell appear valid.
A cell that does not meet all three conditions is not silently treated as a resume. It follows ordinary staging or refusal. The distinction protects the meaning of a pass number and prevents a cell with no prior agent work from inheriting a continuation identity.
4.5 Staging-gate ruling
run_staging_gate begins at run_cell.py line 961 and returns a verdict used by the runner. The normal gate requires the hidden probes to fail on a fresh start tree. The comments explain the reason: a feature cell must lack its feature, or a seeded defect cell must exhibit its defect, before agent work is admitted. For a chained step, the gate can instead return satisfied when the predecessor tree already passes and chained_gate_ruling_enabled at lines 946 to 958 is true.
The ruling is enabled by default. FIVEB_CHAINED_GATE_SATISFIED=0 restores the earlier refusal. This switch is read at call time so a test or one batch can change the ruling without editing the source. The result is not an ordinary hidden pass after agent work. It is a distinct reached state that says the next step may begin from the same tree.
4.6 Reached-step recording
When the gate returns satisfied, _satisfied_by_start_tree at run_cell.py lines 1870 to 1940 clears the live marker, builds a scoring summary with agent_ran false, hidden_pass true, and correctness_gate equal to satisfied_by_start_tree, and writes it to scoring_summary.json and scoring_outputs/scoring_summary.json. It then calls reached_step_metrics at lines 1920 to 1931. That function, defined at lines 3069 to 3119, sets passes, spend, lease wait, and seconds to success to zero, marks the cell admissible, and records the reached status and reason.
The design reason is explicit in the comments. The metrics file is not omitted merely because the cost is zero. The driver uses its presence to decide whether scoring is needed. Earlier omission caused an ordinary scorer to run against a reached cell and replace its verdict. _log_reached at lines 1933 to 1937 and 3122 to 3132 writes a REACHED line rather than NOT-RUN, because the reached step has its own records and is a successful chain participant.
4.7 Chained baseline selection
The scorer calls baseline_tree_for_cell in score_cell.py lines 409 to 469 when it needs a baseline. For a later chain step it reads the task declaration, loads the chain, identifies the head, and prefers the head cell’s staged_snapshot. It falls back to the frozen tree recorded in the current manifest and then to the current cell’s snapshot. The comments explain that comparing only with the immediate predecessor would make the predecessor’s public surfaces appear on both sides and hide the surface the current check needs to find.
For a reached cell, reached_step_summary at score_cell.py lines 1236 to 1255 reads the existing summary and returns it only for the reached correctness state. The scorer entry at lines 1258 to 1268 keeps it and exits unless forced. Thus the ordinary scorer is not allowed to turn a zero-agent, already-passing step into an invalid no-transcript cell by default.
4.8 Chain outcome propagation
After a cell returns, _note_chain_outcome at run_batch.py lines 296 to 325 reads score_states.correctness_gate, falling back to overall. It calls step_was_reached in task_sequences.py lines 69 to 79. Only pass and satisfied_by_start_tree allow continuation. Any other verdict records a reason in the in-memory broken-chain map. The next cell then receives the refusal through _chain_start_tree and is logged without execution.
4.9 Finalisation
The ordinary runner finalises its metrics, manifests, and workspace after the pass loop. A reached step uses the same finalisation boundary after writing its zero-cost records. The batch then reaches its closeout chain in a finally path, where downstream aggregation and reports consume the cell records. The design consequence is that a chain break, a reached step, and a resumed cell all leave durable evidence for the same reporting machinery rather than relying on a process-local explanation.
5. The state machine
The code keeps no state field for a chain. The batch driver holds one dictionary, broken, keyed by the chain key that _chain_key derives from the cell (run_batch.py 257 to 293); a chain is either absent from that dictionary, in which case its next step may run, or present with the sentence that explains why every later step is refused (the chained_sequence_stopped: reason written at lines 284 to 292 and 318 to 325). The verdict of a step is not stored by the driver either: _note_chain_outcome (296 to 325) reads score_states.correctness_gate from the step’s scoring_summary.json and asks task_sequences.step_was_reached whether that verdict lets the chain continue. On the cell side the records that exist are the not_run_reason and budget_exhausted members of metrics.json, the completed status that only a reached step writes, and the satisfied_by_start_tree verdict of the staging gate (task_sequences.py 66). The state names in the diagram below (head, running, step_passed, step_reached, step_failed, chain_broken, resuming, refused, halted_environment, completed) are this chapter’s labels for those situations, chosen so the diagram can be read without the code; none of them is a string the harness writes. The diagram is hand-authored from run_batch.py lines 257 to 325 and 365 to 413, run_cell.py lines 1829 to 1940, and task_sequences.py lines 66 to 79, and each transition names the lines it is drawn from.
stateDiagram-v2 [*] --> head: chain head selected head --> running: run_batch.py 257-293 running --> step_passed: scoring correctness_gate=pass, run_batch.py 296-325 running --> step_reached: staging gate satisfied, run_cell.py 1870-1940 running --> step_failed: other correctness verdict, run_batch.py 296-325 step_passed --> running: predecessor workspace exists, run_batch.py 284-293 step_reached --> running: satisfied_by_start_tree accepted, task_sequences.py 69-79 step_failed --> chain_broken: broken map updated, run_batch.py 318-325 chain_broken --> chain_broken: later step refused, run_batch.py 276-277 running --> resuming: resume flag, workspace, transcript, run_cell.py 1829-1836 resuming --> completed: resumed pass loop finalises running --> halted_environment: interruption leaves invalid record halted_environment --> running: fresh staging when passes remain, run_batch.py 386-407 running --> refused: refusal helper, run_cell.py 1859-1868 step_reached --> completed: metrics and manifests written, run_cell.py 1920-1940 completed --> [*] chain_broken --> [*] refused --> [*]
The state diagram is hand-authored because v001 has no chapter 07 model subgraph. The cited node identifiers available in the flow model for the adjacent ordinary lifecycle are cell.gating, cell.satisfied_by_start_tree, driver.reached_check, driver.resume_decision, driver.resume_run_cell, scorer.reached_check, scorer.reached_kept, scorer.summary_written, aggregate.reached_row, and aggregate.skipped. They describe the same boundary but are not presented as a chapter 07 rendering.
6. The sequence of one unit of work
One unit is one cell in one condition and seed. The participants are the batch driver process, the task-library files, the run-set files, the cell runner process, the scorer process, and the optional agent and hidden-check process. The sequence below is hand-authored from run_batch.py lines 981 to 1033, run_cell.py lines 1829 to 1940, score_cell.py lines 409 to 469 and 1236 to 1268, and task_sequences.py lines 127 to 192.
sequenceDiagram participant Driver as run_batch.py process participant Library as task manifests participant RunSet as run-set files participant Runner as run_cell.py process participant Agent as agent process participant Checks as hidden checks participant Scorer as score_cell.py process Driver->>Library: load_chains, task_sequences.py 127-192 Driver->>RunSet: read prior metrics and scoring records Driver->>Driver: _cell_is_terminal, run_batch.py 365-413 Driver->>RunSet: read predecessor workspace alt chain broken or workspace absent Driver->>RunSet: append NOT-RUN execution log else runnable cell Driver->>Runner: invoke with start tree and resume flag Runner->>RunSet: inspect workspace and transcript alt eligible resume Runner->>RunSet: keep workspace, ledger, snapshot, pass number else new or fresh staged run Runner->>RunSet: stage workspace and snapshot end Runner->>Checks: run staging gate alt start tree already satisfies chained step Checks-->>Runner: satisfied Runner->>RunSet: write reached summary and zero-cost metrics Runner->>RunSet: append REACHED execution log else agent work required Runner->>Agent: invoke measured pass loop Agent-->>Runner: edits, transcript, usage Runner->>RunSet: write metrics and manifests end Runner-->>Driver: cell result Driver->>Scorer: invoke scorer when required Scorer->>RunSet: read baseline and scoring summary alt reached summary retained Scorer->>RunSet: preserve satisfied_by_start_tree summary else ordinary scoring Scorer->>Checks: run visible, full, and hidden checks Scorer->>RunSet: write scoring_summary.json end Scorer-->>Driver: verdict Driver->>RunSet: update chain outcome end
The sequence shows why the reached path does not include an agent or a network service. It is a local gate result based on the predecessor-produced tree. The resume path is also a continuation inside the existing cell folder, not a new cell. The scorer is downstream of both paths and must recognize the reached record before attempting ordinary evidence or transcript-dependent scoring.
7. The records
The canonical records for this mechanism are the cell scoring summary, the cell metrics record, the predecessor workspace, the execution log, and the batch outcome. The batch driver owns the chain decision but does not create a separate chain-outcome JSON file. Its broken-chain map is process state, while its refusal explanation becomes an execution-log line and a result entry. The cell runner owns the reached summary and reached metrics shape. The scorer owns ordinary scoring summaries and the preservation of prior summaries. The repair utility owns only the restoration action and its report.
| Record or table | Writer | Fields or contents | Readers |
|---|---|---|---|
| Predecessor workspace | run_cell.py ordinary staging and finalisation, cell execution region; selected by run_batch.py lines 284 to 293 | Product tree handed forward, including predecessor edits and any handoff artefact | run_batch.py lines 284 to 293; run_cell.py staging and chained gate; score_cell.py lines 409 to 469 through the head snapshot and cell workspace |
scoring_summary.json | run_cell.py lines 1884 to 1918 for a reached step; ordinary scorer later in score_cell.py | cell_id, pass list, visible and hidden pass fields, full-suite fields, agent_ran, validity, score_states, invalid reasons, staging gate, explanation | score_cell.py lines 1245 to 1268; run_batch.py lines 310 to 324 and 375 to 413; restore_reached_step.py lines 43 to 45 and 73 to 93; aggregators read the cell summary through their cell loader |
| Preserved prior scoring summary | score_cell.py preservation helper at lines 1225 to 1230 | Copy of the summary being replaced, named scoring_summary.prior-<stamp>.json | restore_reached_step.py lines 34 to 45 and 73 to 93 |
scoring_outputs/scoring_summary.json | run_cell.py lines 1916 to 1918 for reached status; scorer ordinary output path | Scoring output copy of the summary | Interface and report readers of scoring outputs; repair writes it again at restore_reached_step.py lines 87 to 92 |
metrics.json | run_cell.py lines 1920 to 1932 through reached_step_metrics, or ordinary finalisation | Identity, model, condition, task, seed, timestamps, zero or actual passes, budget, spend, status, agent flag, validity, staging gate, per-pass and category totals | run_batch.py lines 375 to 413; aggregators and report generators; restore_reached_step.py lines 94 to 100 |
execution_log.md | run_batch.py lines 227 to 241 for refusal; run_cell.py lines 3122 to 3132 for reached; repair lines 101 to 106 | Timestamp, cell identifier, verb such as NOT-RUN or REACHED, and reason | Interface and report readers; the batch log is also the durable explanation of a broken chain |
| Run-set result list | run_batch.py lines 981 to 1033 | Cell identifier, skipped reason, or returned cell result | Batch closeout and batch_outcome.json construction |
batch_outcome.json | run_batch.py closeout region after the cell loop | Run identifier, status, reason, planned, attempted, and skipped counts | Run-set reports, monitoring, and interface readers |
| Campaign metric row | Aggregation after the batch, not owned by this chapter | Reached row or ordinary cell outcome derived from metrics.json and scoring_summary.json | Campaign tables and reports; the flow model names aggregate.reached_row and aggregate.skipped |
The reached summary and metrics are deliberately separate. The summary is the correctness verdict consumed by the scorer and chain driver. Metrics are the cost and identity record consumed by the driver and aggregators. A zero-cost step still needs both, because absence of metrics previously caused the driver to invoke ordinary scoring and lose the reached verdict.
flowchart LR M[task_manifest.json] --> L[task_sequences.load_chains] L --> D[run_batch.py chain decision] D --> W[predecessor workspace] W --> G[run_cell.py staging gate] G --> S[scoring_summary.json] G --> X[metrics.json] G --> E[execution_log.md] S --> C[score_cell.py reached check] X --> C S --> B[run_batch.py chain outcome] B --> O[batch_outcome.json] X --> A[aggregate reached row] S --> A O --> T[campaign tables] A --> T P[scoring_summary.prior-stamp.json] --> R[restore_reached_step.py] R --> S R --> X R --> E
This lineage diagram is hand-authored from the writer and reader locations in the table. It is a record-lineage diagram from the mechanism’s files to campaign tables, not a rendering of the missing chapter 07 flow-model subgraph.
8. The loops and the waits
The task-library loader has a bounded manifest loop over sorted paths in task_sequences.py lines 137 to 146. Its bound is the number of matching task manifests, and it stops after the glob is exhausted. The backward and forward chain walks in load_chains are bounded by the number of declared members when declarations are valid. A circular declaration raises SequenceDeclarationError; a branch with more than one successor raises the same error. The validation function loops over grouped chain entries and stops at the end of the batch specification. The pruning function advances its reach while the next chain member is present and stops at the first absent member or the chain length.
The batch cell loop at run_batch.py lines 981 to 1040 is bounded by spec["cells"] and stops at end, or earlier when the batch aborts. For each cell it can invoke the runner and scorer once for the initial attempt, then the surrounding resume logic can invoke further passes while a hidden failure remains, passes remain, and the cumulative token budget permits it. The exact resume-loop guard is in the driver flow outside the chain predicates, while the cell-level eligibility predicate is run_batch.py lines 365 to 413.
The runner’s pass loop is bounded by max_passes and the token budget. It stops on a successful hidden result, a refusal, an invocation or protocol error, an exhausted budget, a halted environment, or spent passes. A resume does not reset either bound. Its stop condition is therefore cumulative across the original and resumed invocations. The scorer’s visible flakiness rerun is bounded by its fixed three-way check in the ordinary scoring path. Reached scoring does not enter that loop.
There is no wait in the chain selection, reached-step, or restoration paths. They perform local file reads and writes. A resumed ordinary cell may wait inside the agent backend for a lease or service, but that interval belongs to the backend lifecycle and is not created by the chain mechanism. The chain mechanism has no retry policy for a missing predecessor workspace: it refuses and records the reason. Restoration likewise does not retry a missing prior summary. It reports that there is nothing to restore.
9. The guards and refusals
| Guard | Location | Effect and refused condition |
|---|---|---|
| Declaration readability | task_sequences.py lines 137 to 146 | Skips an unreadable task manifest rather than making it a usable dependency. |
| Circular declaration | task_sequences.py lines 157 to 167 and 185 to 188 | Raises SequenceDeclarationError; a loop cannot form a linear chain. |
| Branch declaration | task_sequences.py lines 170 to 183 | Raises SequenceDeclarationError; one produced tree cannot be carried into two successors. |
| Missing chain prefix | task_sequences.py lines 239 to 255 and prune_orphan_steps lines 340 to 361 | Reports or removes a later task whose earlier task is absent for the same condition and seed. |
| Wrong chain order | task_sequences.py lines 257 to 269 | Reports a batch that runs a later step before its predecessor. |
| Scattered chain | task_sequences.py lines 271 to 287 | Reports a chain interrupted by unrelated cells. |
| Broken-chain map | run_batch.py lines 273 to 277 | Refuses every later step after the chain has been marked broken. |
| Missing predecessor workspace | run_batch.py lines 284 to 293 | Refuses a later step rather than silently using the frozen build. |
| Predecessor correctness | run_batch.py lines 296 to 325 and task_sequences.py lines 69 to 79 | Continues only for pass or satisfied_by_start_tree; every other correctness state breaks the chain. |
| Terminal cell | run_batch.py lines 365 to 413 | Skips an existing cell whose verdict, budget, scorer state, or pass count says no further run is allowed. |
| Halted-cell exception | run_batch.py lines 386 to 407 | Prevents a halted cell with an attempt remaining from being skipped as terminal. |
| Resume eligibility | run_cell.py lines 1829 to 1836 | Does not resume unless the flag, workspace, and streamed transcript are all present. |
| Sticky invalidity | run_cell.py lines 1838 to 1857 | Refuses to erase an earlier invalid-for-analysis result during continuation. |
| Chained gate ruling | run_cell.py lines 946 to 958 | Uses the reached state unless the environment is exactly 0, in which case the old refusal is restored. |
| Staging gate | run_cell.py lines 961 to 980 and the reached call site at 1870 | Refuses ordinary measured work when the start tree already passes; for an enabled chained step it changes the result to reached. |
| Reached summary recognition | score_cell.py lines 1236 to 1268 | Keeps an existing reached summary and refuses ordinary rescoring unless forced. |
| Restore prior selection | restore_reached_step.py lines 34 to 45 and 75 to 83 | Does nothing when neither the current nor a preserved summary records the reached state. |
| Cell identity shape | restore_reached_step.py lines 48 to 53 | Exits when the target directory is not shaped as a four-part cell folder. |
These guards divide responsibility. Sequence declarations protect the experimental plan. The driver protects source-tree continuity. The runner protects the meaning of admission and continuation. The scorer protects the reached verdict from an incompatible ordinary path. The repair utility refuses to invent a prior verdict.
10. The unhappy paths
Every path below is described in four parts. The first part states what the step is supposed to do. The second explains why the implementation works that way. The third identifies the trigger and the record left. The fourth states the cost to the driver, scorer, aggregator, and interface.
| Trigger and four-part account | Code path and record | Status and exit code | Downstream effect |
|---|---|---|---|
A later chain step has no predecessor workspace. The step is supposed to begin from the predecessor’s produced tree. The driver works this way to prevent silent fallback to the frozen build. The missing directory triggers refusal and leaves a NOT-RUN line plus a skipped result. | run_batch.py lines 284 to 293 and 1013 to 1023; execution_log.md is the durable record. | The cell is skipped; the batch continues, with no child exit code because no child is invoked. | The driver refuses later steps in the same chain. The scorer is not invoked. The aggregator sees a skipped cell. The interface can display the logged reason. |
A predecessor has a non-passing correctness state. The next step is supposed to consume only a tree in which the earlier feature is known to work. The map treats only pass and reached as safe. The failed or absent verdict triggers chained_sequence_stopped and leaves a broken in-memory chain plus a NOT-RUN line on the next refusal. | run_batch.py lines 296 to 325, 1016 to 1023; task_sequences.py lines 69 to 79. | The predecessor retains its scorer status; later cells are not run. | The driver skips descendants. The scorer records only the predecessor’s result. Aggregation preserves the failure and skipped descendants. The interface shows the refusal reason. |
| The batch specification has a gap, wrong order, or scattered chain. The plan is supposed to run a chain head-first and consecutively. Validation works this way because a later task from the wrong tree has a different measurement meaning. The declaration triggers a violation before execution, leaving the validation messages and no valid cell execution. | task_sequences.py lines 222 to 289; optional correction through lines 292 to 364. | The preflight or plan validation refuses or prunes according to its caller; the validation function itself returns violations rather than a process status. | The driver should not create measured cells for the invalid arrangement. The scorer and aggregator receive no new chain records. The interface can show the complete sentence returned by validation. |
The start tree already passes a chained step. The step is supposed to continue the chain without spending agent work. The distinct reached state works this way so a genuine zero-cost success is not confused with a skipped cell. The staging gate triggers satisfied, and the runner writes the scoring summary, metrics, scoring output, and REACHED log line. | run_cell.py lines 1870 to 1940 and 3069 to 3119. | Status is satisfied_by_start_tree; the reached path returns normally with zero passes and zero tokens. | The driver can continue the chain. The scorer keeps the summary. Aggregation can emit a reached row. The interface can show a completed zero-cost step rather than a skip. |
| The reached metrics record is absent or an old scorer overwrites the reached summary. The step is supposed to remain an admissible reached success. The metrics file is required because the driver uses it to decide whether scoring is needed, and the prior copy is required because the scorer preserves replaced verdicts. The missing or overwritten record triggers repair. | restore_reached_step.py lines 34 to 107; it restores the summary, output copy, metrics, and log where applicable. | Normal repair returns JSON with actions and exit code 0. With no prior reached summary it returns a no-action report and exit code 0. | The driver can recognize the finished step after repair. The scorer sees the reached state. Aggregation sees zero-cost metrics. The interface sees the repair log. |
| Resume is requested without all eligibility evidence. The cell is supposed to continue only the same measured attempt. The three-part predicate prevents a new or empty folder from being presented as continuation. The missing flag, workspace, or transcript triggers ordinary non-resume handling. | run_cell.py lines 1829 to 1836; ordinary staging or refusal writes the normal cell records. | No special resume status and no separate exit code. The caller receives the ordinary cell result. | The driver either starts a valid cell or records its ordinary refusal. The scorer follows the resulting records. Aggregation does not merge an unqualified attempt into a prior cell. The interface sees the actual status. |
| A resumed cell was previously invalid. The step is supposed to preserve the analysis qualification of the measured cell. The invalidity is sticky because rebuilding the result object otherwise erases the only evidence of the earlier guard. The prior metrics trigger carry-forward fields. | run_cell.py lines 1838 to 1857; updated metrics.json at finalisation. | The cell may continue, but remains invalid for primary analysis. | The driver can still use its pass and budget state. The scorer may produce a correctness verdict, but aggregation excludes or marks the row invalid. The interface can show the carried reason. |
| A cell is halted part way through. The step is supposed to be rerun from a safe staged tree, not resumed from partial edits. The halt predicate distinguishes interrupted invalidity from intentional validation. A halt with an attempt remaining makes the cell non-terminal. | run_batch.py lines 327 to 363 and 386 to 407; the scorer may already have written a summary for the half-finished tree. | The next resume-into walk does not skip it; a fresh run is selected. No special child code is assigned by this predicate. | The driver spends a new allowed attempt when available. The scorer later scores the fresh result. Aggregation can retain the earlier invalid evidence and the replacement. The interface avoids reporting a permanently resumable or completed halt. |
| The cell exhausts passes or budget, or ordinary checks fail. The step is supposed to stop under the declared protocol. The terminal predicate treats hidden pass, missing scorer, visible failure, budget exhaustion, and spent passes as terminal. The trigger leaves ordinary metrics and scoring summary. | run_batch.py lines 365 to 413 and ordinary finalisation in run_cell.py. | The cell remains failed, exhausted, or otherwise terminal according to its records; the child exit code is the ordinary runner result. | The driver does not resume it. The scorer’s verdict feeds aggregation. The interface reports the terminal status and cost. |
| The cell folder has the wrong identity shape during repair. The repair is supposed to operate on one known cell. The parser refuses malformed identity rather than guessing project, task, phase, or seed. The trigger leaves no repair record. | restore_reached_step.py lines 48 to 53. | SystemExit from _identity; the command exits nonzero. | The driver is unaffected unless an operator invoked repair as a required step. The scorer and aggregator see no changed record. The interface sees no new log line. |
A path that writes nothing is material here. The malformed repair identity and a no-op restore do not invent metrics or verdicts. A chain refusal writes only an explanation, not a fake scoring summary. This distinction keeps the aggregator from counting a cell that never had a measured tree.
11. The metrics produced or fed
The reached path produces a normal-shaped metrics record with zero activity rather than omitting the row. reached_step_metrics at run_cell.py lines 3069 to 3119 maps the identity inputs to cell_id, run_id, project_id, condition, task_id, phase_id, seed, model fields, timestamps, and task directory. It maps the outcome to seconds_to_success: 0.0, lease_wait_seconds: 0.0, passes_run: 0, budget_exhausted: false, cumulative_spend: 0, agent_invocation_error: false, invalid_for_primary_analysis: false, status: satisfied_by_start_tree, agent_ran: false, and not_run_reason: start_tree_satisfies_step. The staging gate and stage timing are retained as advisory provenance, while per_pass and category_totals are empty because no pass existed.
The scoring summary maps the reached state to hidden_pass: true, visible_pass: null, full_suite_pass: null, agent_ran: false, valid_for_primary_analysis: true, and score_states.correctness_gate: satisfied_by_start_tree. The null visible and full-suite fields mean those checks were not run by an agent pass, not that they failed. The explanation states that the predecessor tree passed the hidden checks, no tokens were spent, and the next task continues from the unchanged tree.
For an ordinary chained step, baseline_tree_for_cell maps the chain head’s staged snapshot, or its recorded frozen tree, into the scorer’s baseline input. The baseline choice is advisory to the scorer’s comparison logic but substantive to correctness: it determines which inherited surfaces are treated as the chain’s starting material. The scorer maps its verdict to score_states, pass fields, invalid reasons, and the summary read by the driver. The driver maps pass or satisfied_by_start_tree to chain continuation and every other correctness state to a broken chain.
At closeout, aggregators map reached metrics to a campaign row through the aggregate reached-row path named in the flow model as aggregate.reached_row. A reached row must remain distinguishable from an ordinary skipped row. The execution log is advisory narrative for the interface and reports, while metrics.json and scoring_summary.json are the analytic inputs. No network usage, lease duration, token usage, or agent pass is inferred for a reached step; each remains zero or null as written.
12. The tests
The directly relevant test files are 5. Experiment/1. Harness/scripts/tests/test_task_sequences.py, test_chained_gate_satisfied.py, test_reached_step_scoring.py, test_chained_baseline.py, test_run_batch_sequences.py, test_resume_into.py, and test_run_batch.py. The isolation tests also exercise the staging contract that supplies the gate. The test files are named here because they are present in the harness tree and their module descriptions and references identify the mechanism under test.
test_task_sequences.py exercises chain declaration, the reached verdict, ordering, grouping, and orphan pruning. test_chained_gate_satisfied.py exercises the operator ruling, the reached record shape, and the interaction between run_cell.py and score_cell.py. test_reached_step_scoring.py exercises the case in which a reached cell is sent through the scorer and the existing reached summary must be retained. test_chained_baseline.py exercises selection of the chain-head baseline rather than the immediate predecessor snapshot. test_run_batch_sequences.py exercises ordered chain execution, refusal after a predecessor failure, and the driver’s skipped-cell records. test_resume_into.py exercises existing run-set continuation and the arguments passed to the cell runner. test_run_batch.py exercises the batch driver invocation and resume behavior. test_isolation_contract.py contains the staging and hidden-check contract that makes an already-satisfied start tree a meaningful gate result.
The mapping of tests to every exit path in section 10 was not made. The sources identify relevant test files, but they do not establish a one-test-per-path coverage matrix for missing workspaces, malformed repair identities, all terminal predicate branches, every no-op restore case, or each combination of resume evidence. The statement is therefore not that every unhappy path is covered. It is that the named test files exercise the central chain, reached, baseline, scoring, and resume mechanisms, while the complete exit-path mapping remains unchecked.
13. The dated incidents that shaped the code
The comments in task_sequences.py say that on 2026-08-30 twenty-one cells were run from a frozen build after a chain prefix was omitted, and one passing cell had first rebuilt its predecessor’s entire feature by adding 2,761 lines. This is the reason the declaration loader and validator require a head-first, consecutive chain for one condition and seed. The historical account appears in task_sequences.py lines 17 to 26, and the resulting rule is stated at lines 28 to 37.
The staging-gate comment at run_cell.py lines 965 to 970 records the 2026-08-08 run set 007 re-score finding. Three feature cells were staged with the feature already built, so hidden probes scored the reference implementation rather than agent work. The gate was added to refuse measured admission when the start tree already passed. The chained exception is explained at run_cell.py lines 978 to 980 and the ruling at lines 950 to 956: a chained step that is already fulfilled must be marked reached so future sequence steps remain reachable.
The sticky-invalidity comment at run_cell.py lines 1838 to 1844 records run set 009 and cell H-LAP-L5P__H-M018__full__s2. An escape guard fired at pass 2, but a later resume finalisation overwrote the invalid result with completed and valid. The carry-forward fields at lines 1845 to 1857 exist so a resumed cell cannot erase that analysis warning.
The reached metrics comment at run_cell.py lines 3078 to 3087 records run set 073 and step H-M035 on 2026-09-03. The earlier reached implementation wrote a scoring summary but no metrics. The driver therefore ran the ordinary scorer on a cell with no transcript and no manifest, and the scorer overwrote the reached verdict as invalid. The current zero-count metrics record and the scorer’s reached check are the repair represented by test_chained_gate_satisfied.py and test_reached_step_scoring.py.
The restoration utility’s module comment at restore_reached_step.py lines 4 to 15 gives the same 2026-09-03 incident and names tracker 2.35. It records that the scorer preserves the replaced verdict as scoring_summary.prior-<stamp>.json, which lets the repair utility restore the summary, write the metrics shape, and append an explanation rather than reconstructing a verdict from guesswork.
The baseline comment at score_cell.py lines 418 to 426 records five cells failing between run sets 055 and 070 because a chained step was compared with the immediate predecessor tree. The repair at lines 428 to 436 makes the chain head’s starting tree the preferred baseline, so inherited surfaces remain visible to the hidden comparison as work already established by the chain.
The comments are design history, not evidence that these guards are accidental. Each incident explains a refusal, a preserved field, or a separate reached state that would otherwise look unnecessarily strict.
14. The weakest claim, what was not checked, and the token line
The weakest claim is that every downstream report and interface consumer will interpret a repaired or reached record identically to a record produced in the original cell run. The cited code verifies the writer, the chain reader, the scorer guard, and the repair actions, but it does not verify every campaign-table schema, every interface branch, or a complete test mapping for every unhappy path. The flow model v001 has no chapter 07 subgraph, so the diagrams here are source-derived and hand-authored rather than regenerated from that model. The mapping of tests to exit paths was not made. These are the unchecked boundaries.
Model: openai/gpt-5.6-luna via OpenRouter; tokens: see the job ledger