Written 2026-09-12 by Claude Fable 5.1 to the standard set in 5. Experiment/11. Detailed Design/specifications/C0-detailed-design-specification.md and shown by the reviewed sample chapter 06-execution-of-a-cell.md. Every line number below was read from the source at commit 6e2ee2c2b; the scripts this chapter owns and the batch driver have the same Git blob hashes at that commit as at commit b69ae5977, the commit the flow model 5. Experiment/11. Detailed Design/flow-model/flow_model.v001.json was verified against. The local-model draft operations/detailed-design/drafts/chapter_10.md (100,932 bytes) was the lead for the file inventory and nothing else; section 14.1 lists what it got wrong. The flow model’s chapter 10 subgraph (eighteen nodes, thirty-seven edges of which one is inferred, no declared loops) is the source of the first state diagram of section 5, and every node identifier is named so the renderer can regenerate it.
Owning files. Fourteen scripts under 5. Experiment/1. Harness/scripts/: aggregate_metrics.py (1,202 lines, 55,029 bytes), aggregate_timings.py (1,188 lines, 56,864), update_monitoring.py (251 lines, 9,830), gen_iteration_summary.py (503 lines, 22,182), gen_run_set_report.py (2,222 lines, 104,448), gen_run_set_analysis.py (836 lines, 38,178), gen_cell_narrative.py (509 lines, 22,866), cleanup_run_set.py (254 lines, 9,649), publish_experiment_results.py (299 lines, 12,298), backfill_closeout.py (158 lines, 6,133), metrics_db.py (973 lines, 44,171), extract_cell_metrics.py (1,315 lines, 55,380), reflow_markdown.py (159 lines, 4,624) and lint_report_prose.py (160 lines, 4,963); and the normative document 5. Experiment/6. Metrics/metric_definitions.md (32,835 bytes). Files this chapter reads into but does not own: run_batch.py lines 543 to 599 and 600 to 895 (the closeout chain and its progress record, inside the batch driver that chapter 02 owns), publish_lock.py (the campaign-wide lock, chapter 02), timing_lib.py and ledger_lib.py (chapter 09), run_defaults.json (chapter 01), 10. User Interface/server.py and auto_experiment.py (chapter 11), publish_snapshot.py and paired_analysis.py (chapter 12), and attest_batch_outcome.py (chapter 02).
1. What the closeout is and why it runs the way it does
A cell is one measured attempt: one coding agent, one maintenance task, one product build, one repetition seed, run at most five times within one cumulative token budget. A run set is one batch of cells, and the closeout is what happens when the batch ends: seven programs, run in a fixed order by the batch driver, turn the cells’ own records into the campaign tables, refresh the monitoring files, write the run set’s research article and its results report, draft the written analysis, delete the regenerable files the cells left behind, and commit the results to the research repository and the public portal. The chain is declared in one table in the driver, CLOSEOUT_CHAIN (run_batch.py lines 564 to 595), whose comment states the two properties that matter: each stage has a stable key so a reader can follow it across rewrites of the record, and the order matters because each program reads files an earlier one wrote (558 to 563).
The chain runs in the driver’s finally block, so it runs however the batch ended: after the last cell, after the circuit breaker, after a stop signal, or after a crash (1046 to 1064). The docstring records why (17 to 24): before 2026-08-31 the closeout sat after the cell loop inside the try, a stop signal raised straight through it, and run set 041 ran eleven cells with real verdicts, was stopped on purpose, and left no trace in the campaign tables; the runs an operator stops are disproportionately the ones that found something, so losing exactly those is the worst selection bias an evidence log can have. The whole chain holds the campaign-wide publish lock (833 to 849; publish_lock.py), because two endpoint lanes finishing at once would otherwise overwrite each other’s rows in the files the campaign owns.
Every stage is guarded on its own: a stage that fails is recorded as failed and the next stage runs anyway (871 to 875, 883 to 888), and a stage that prints a line beginning SKIPPED: is recorded as skipped with its reason (877 to 880). The progress record closeout_progress.json is rewritten around every stage so a reader, and the results interface, can see what is pending, running, done, skipped or failed while the chain is still running (685 to 719). The record therefore never says a closeout completed when a stage failed: the closeout’s own status is failed when any stage was (893 to 895), and a failed publication can be replayed later by backfill_closeout.py without re-running the six stages before it.
The first stage, the aggregator, is where a cell becomes a row. Its gate (aggregate_metrics.py 1098 to 1141) admits a cell only when it has a metrics record, a token ledger and a scoring verdict, or when it is a chained step reached without an agent; everything the results-integrity review of 2026-09-12 found about which cells are counted (operations/results-integrity/B2-count-provenance.md, section 1) passes through this gate. The tables it writes under 6. Metrics/ are cumulative across the campaign and keyed by run set and cell, so re-aggregating a run set rewrites its rows in place; the two JSON files beside them are rewritten wholesale on every invocation and describe only the last aggregation’s cells (6. Metrics/README.md, the paragraph beginning “Two properties of the writer matter”).
Four of the fourteen owning scripts are not stages. aggregate_timings.py builds the campaign timing tables and is run by the results interface on its first request after a run set changes, not by the closeout (server.py 3247 to 3278; the flow model’s open item 16). gen_cell_narrative.py writes a chronological account of one cell by hand. extract_cell_metrics.py and metrics_db.py are the offline behavioural extraction and its database, which the report generator calls into by name when they are importable (gen_run_set_report.py 1058 to 1073). reflow_markdown.py and lint_report_prose.py are the house style’s tool and its check: every generated document is unwrapped to one physical line per paragraph before it is written, and the lint fails a report that wraps its prose.
2. The reader’s map of the owning files
2.1 The chain in the driver, run_batch.py lines 543 to 895
_write_active_batch (543 to 555) rewrites the campaign’s one-line statement of which batch is running or last finished. CLOSEOUT_CHAIN (564 to 595) names the seven stages: key, title, explanation, script, argument builder. CLOSEOUT_SCRIPT_NAMES (597) is the set the invoker captures output for, and SKIPPED_PREFIX (598) the sentence a program prints to decline its work. _stop_reason (601 to 614) puts a stop into plain words. _write_batch_outcome (617 to 650) writes the run set’s account of how the batch ended. _new_closeout_progress (658 to 682) builds the record with every stage pending; _write_closeout_progress (685 to 719) replaces it atomically. _invocation_parts (722 to 729), _last_output_line (732 to 738), _skipped_reason (741 to 748) and _printed_artifacts (751 to 770) read a stage’s exit status and output. set_run_index_status (773 to 830) closes the run set’s row in the campaign index. _closeout (833 to 849) takes the lock and _closeout_locked (852 to 895) runs the chain. The invoker default_invoke (113 to 140) captures a closeout program’s output and repeats it to the launcher’s log (its docstring, 114 to 122).
2.2 aggregate_metrics.py
Seven regions. Lines 1 to 39 are the docstring and the imports; the docstring lists the cell files it reads and the five tables it writes. Lines 41 to 105 are the paths, the budget and grid constants (46 to 47), the factor columns (48), the three column-addition tuples for the matrix (TIMING_ADD_COLUMNS 56, NEW_CONTEXT_ADD_COLUMNS 83, HANDOFF_ADD_COLUMNS 91), the reconciliation tolerance (94 to 95) and the artefact kinds (98). Lines 107 to 282 are the readers: _load_defaults (107), the ledger helpers (120 to 149), load_snapshot (152), payload_tokens (165), cone_files (183), _read_stream (197), parse_transcript (206), _pass_transcripts (223), reconcile_check (244 to 267) and _reported_usage (270). Lines 284 to 359 parse a patch and decide the success pass. Lines 361 to 690 load one cell (load_cell 361, _load_task 373, _context_window 387) and compute its row (compute_cell 398 to 687). Lines 691 to 810 are the metric helpers. Lines 811 to 1032 write and pair: _read_csv (811), upsert_csv (819 to 860), _fmt (863), build_pairings (877 to 900), _pairing_row (903 to 946), pass_at_budget (955), group_analysis (975) and the dispersion helpers. Lines 1034 to 1202 are the run-set loop: _iter_cell_dirs (1034), reached_cell_metrics (1046 to 1095), process_run_set (1098 to 1141), write_all (1144 to 1168), _surface_row (1171) and main (1183 to 1198).
2.3 aggregate_timings.py
Its docstring (1 to 46) names the three files it produces and the rule that a measured cell and a reconstructed one are never mixed silently. Lines 65 to 219 are the paths, the reconstruction caps (94 to 112), the column maps (114 to 175) and the two column lists (177 to 219). Lines 222 to 305 are helpers. reconstruct (307 to 460) rebuilds a cell’s activities from its artefacts when it has no timing file; model_loads_during (466 to 486) attributes a model start-up to a waiting cell; cell_rows (489 to 655) builds one cell’s wide row and long rows; _fill_request_columns (658), _working_span (712), _cell_end (754), summarize_attempts (771), _counts_towards_results (853 to 865) and summarize_campaign (868 to 968) roll up; _run_sets (974), _cell_order (987), place_derived_rows (1007 to 1051), collect (1054 to 1114), _write_csv (1117), _cell_value (1126) and main (1137 to 1184) drive.
2.4 The five smaller stages and the two publishers
update_monitoring.py: the proxy threshold for the ledger warning (31), the five table names (33 to 39), update_monitoring (70 to 112), _write_monitor_state (115 to 135), _write_live_summary (138 to 173), _append_run_history (176 to 198), _write_dashboard (201 to 241) and main (244 to 247). gen_iteration_summary.py: the docstring recording the 2026-09-02 merge (1 to 37), _rows (63 to 79), the SVG figure functions (88 to 202), _draw_figures (204 to 237), build_appendix (240 to 250), the article skeleton (254 to 349), _figure_block (351), _design_lines (364 to 394), ensure_article (397 to 421), reset_draft_article (424 to 449), refresh_appendix (460 to 479), update_index (482 to 491) and the entry point (494 to 503). gen_run_set_report.py: the docstring with the two locations and the two-part rule (1 to 42), the constants including the two part markers, their legacy wordings, the part-two notice, the draft marker and the five part-two sections (63 to 156), the readers (159 to 224), the formatting helpers (226 to 306), the eighteen section builders (308 to 1645), _part_two (1647), _numbered (1660), build_full_report (1675 to 1736), build_summary (1739 to 1815), the article appendix builders (1817 to 2009), _find_part_two (2020), part_two_state (2035 to 2054), _merge_preserving_part_two (2057 to 2076), refresh_summary (2079 to 2097), write_report (2100 to 2131), _index_summary (2138 to 2185) and main (2195 to 2218); the extractor hooks and their time limits are at 1026 to 1221. gen_run_set_analysis.py: the docstring stating why it exists, what it never does, and how a draft is told from a judgment (1 to 68), the settings (84 to 124), drafting_enabled, analyst_model and analyst_executable (136 to 163), the system prompt (178), ask_model (220 to 243), _run_subprocess (245 to 259), the house-style repairs (262 to 321), the prompts and checks (324 to 524), _data_of_record (526), draft_run_set (547 to 672), _draft_article (675 to 739), _cells_in_earlier_draft (742), _write_record (758) and main (775 to 832). cleanup_run_set.py: the deletable names (29 to 41), the two allowed parents (43), CleanupRefused (47), the five checks and finders (51 to 152), _write_cleanup_record (155 to 176), cleanup_run_set (179 to 226) and main (229 to 250). publish_experiment_results.py: the paths and the trailer (29 to 38), PublishError (41), the Git helpers (45 to 61), the five preconditions (64 to 111), _commit_paths (114 to 129), _build_portal_candidate (132 to 145), _publish_portal_snapshot (148 to 187), _read_outcome (190 to 201), _source_paths (204 to 218), publish (221 to 263), _enabled (266) and main (271 to 295). backfill_closeout.py: the progress helpers (36 to 69), _record_recovery (72 to 100), backfill (103 to 130) and main (133 to 154).
2.5 The four tools outside the chain
gen_cell_narrative.py: the retry notices mirrored from the runner (28 to 45), the parsers (53 to 152), narrate (154 to 469) and main (472 to 505). extract_cell_metrics.py: the docstring stating that every number is a deterministic function of preserved files (1 to 27), the truncation markers and tool sets (56 to 82), the parsers (84 to 362), the file composition measures (362 to 567), the identity and root helpers (569 to 716), extract_cell (718 to 1031), _rollup (1033), _doc_artifact_rows (1219), iter_cell_dirs (1267) and main (1278 to 1311). metrics_db.py: the docstring with the four questions the database answers (1 to 37), the schema (59 to 356), the connection and insertion helpers (358 to 464), run_number, list_run_sets, already_extracted and backfill (466 to 540), the factor-matrix mirror (547 to 606), crosscheck_artifact_use (613), the exports (701 to 741), run_set_report_metrics (774 to 879) and main (881 to 969). reflow_markdown.py: the seven patterns (12 to 19), reflow (22 to 143) and main (146 to 155). lint_report_prose.py: the write-up names (22 to 28), _prose_lines (41 to 83), hard_wraps (86 to 107), report_paths (110 to 124), scan (127 to 133) and main (136 to 156).
2.6 metric_definitions.md
The normative definition of every metric: the correctness gate (section 1), the primary metric (2), the docs-subtracted metrics (3), artifact_use_evidence (4), the other secondary metrics (5), validity (6), the CSV schema of the five tables with the three additive updates of 2026-07-11, 2026-07-17 and 2026-09-01 to 2026-09-02 (7, lines 144 to 231), the generation-phase metrics (8) and the thirteen pre-registered additions (9). Its section 7 carries a migration note of 2026-08-11 that the aggregator still computes the paired delta against the retired best-baseline selection, which the code confirms (aggregate_metrics.py 877 to 946).
3. The inputs
3.1 Command-line arguments
| Script | Parser | Arguments |
|---|---|---|
aggregate_metrics.py | 1184 to 1186 | --run-set (one folder; absent means every run set with a cells/ folder, 1187 to 1192). The closeout passes none (run_batch.py 568) |
aggregate_timings.py | 1138 to 1148 | --base, --run-set (repeatable, prefix match), --out (default 6. Metrics) |
update_monitoring.py | 245 | none |
gen_iteration_summary.py | 495 to 497 | --run-id (required); the closeout passes the run set’s name (579) |
gen_run_set_report.py | 2196 to 2205 | --run-set or --run-id (one required), --base; the closeout passes --run-set (583) |
gen_run_set_analysis.py | 776 to 787 | --run-set or --run-id, --base, --redraft, --dry-run, --model; the closeout passes --run-set (588) |
cleanup_run_set.py | 230 to 236 | --run-set (required), --apply; the closeout passes both (592) |
publish_experiment_results.py | 272 to 284 | --run-set (required, the folder name), --portal, --remote, --branch; the closeout passes --run-set (595) |
backfill_closeout.py | 134 to 145 | --run-set (repeatable; absent means every run set whose publication failed, 146), --portal, --remote, --branch, --dry-run |
gen_cell_narrative.py | 473 to 479 | --run-set (required), --cell or --all, --stdout |
extract_cell_metrics.py | 1279 to 1285 | a cell folder or --run-set, --json, --with-tree-composition |
metrics_db.py | 882 to 907 | --db, --init, --backfill, --from, --to, --run-set (repeatable), --import-factor-matrix, --crosscheck, --export-csv, --resume, --with-tree-composition, --verify-roundtrip, --regen-existing |
reflow_markdown.py | 146 to 147 | file paths, rewritten in place |
lint_report_prose.py | 137 to 140 | --root |
3.2 Environment variables
| Variable | Read at | What it decides |
|---|---|---|
FIVEB_AUTO_PUBLISH | publish_experiment_results._enabled 266 to 268 | 0, false, no or off makes the publication stage print SKIPPED: and exit 0 (285 to 288), for an offline or private run (run_batch.py 589 to 591) |
PHD_PORTAL_ROOT, FIVEB_PUBLISH_REMOTE, FIVEB_PUBLISH_BRANCH | publish_experiment_results.main 276 to 283; backfill_closeout.main 138 to 140 | The portal checkout (default /media/david/Windows/SourceCode/infinitebydesign/phd, 32), the remote and the branch (defaults origin and main) |
FIVEB_ANALYSIS_DRAFT | gen_run_set_analysis.drafting_enabled 136 to 139 (the name at 117) | Off makes the drafting stage print SKIPPED: and exit 0 (796 to 799) |
FIVEB_ANALYST_MODEL, FIVEB_ANALYST_EXECUTABLE | analyst_model 141 to 143, analyst_executable 145 to 163 (names at 93 and 95) | The analyst model, default claude-opus-5 (94), and its command-line tool, default claude (96), searched in known places when not on the path (108 to 111) |
FIVEB_ANALYSIS_TIMEOUT_SECONDS | _env_float 127 to 134 at 559 to 560 (the name at 113) | How long one model call may take, default 900 (114) |
FIVEB_REPORT_EXTRACTOR_SECONDS, FIVEB_REPORT_EXTRACTOR_CELL_SECONDS | gen_run_set_report._env_seconds 86 to 91 at 1180 to 1183 | The transcript reader’s budget for a whole run set (default 180, 82) and for one cell (default 30, 83) |
No other owning script reads an environment variable. The drafter passes the model’s own command line a fixed argument set with every tool switched off (ask_model 230 to 232).
3.3 Files read
| File | Read at | What it decides |
|---|---|---|
cells/<cell>/metrics.json | aggregate_metrics.process_run_set 1103 (the reached-step preview) and load_cell 364; gen_run_set_report.load_cell 176; cleanup_run_set._finished_without_a_score 91; aggregate_timings.cell_rows 495; gen_cell_narrative.narrate 156; extract_cell_metrics 12 | The cell’s identity, status, spend, timing roll-up and flags (chapter 06, record table) |
cells/<cell>/scoring_summary.json | aggregate_metrics.load_cell 365 and _success_pass 353; gen_run_set_report.load_cell 177; cleanup_run_set._require_scored_cells 108; aggregate_timings.cell_rows 497; gen_cell_narrative.narrate 157 | The verdict of record and the success pass (chapter 08) |
cells/<cell>/token_ledger.jsonl | aggregate_metrics.load_cell 366 and the gate at 1113 | Every token column (chapter 09) |
cells/<cell>/transcripts/pass<k>.stream.jsonl | aggregate_metrics._pass_transcripts 223 (through parse_transcript 206: usage, result, compactions); gen_run_set_report through the extractor hook 1153; extract_cell_metrics.parse_transcript 243; aggregate_timings through timing_lib | Reconciliation, cost, context health, the navigation measures |
cells/<cell>/patches/pass<k>.patch | aggregate_metrics._load_patch 770; extract_cell_metrics.parse_patch 331; gen_cell_narrative.parse_patch 96 | Diff statistics and outside-cone edits |
cells/<cell>/lap_path_manifest_snapshot.json | aggregate_metrics.load_snapshot 152; extract_cell_metrics 748 | The documentation artefact set and the payload denominator |
cells/<cell>/cell_manifest.json | aggregate_metrics.compute_cell 516 to 523; gen_run_set_report.load_cell 178; aggregate_timings.cell_rows 496; gen_cell_narrative.narrate 155 | The registered test command and shell_available; the container and lease records |
cells/<cell>/hidden_tier_result.json | gen_run_set_report.load_cell 179 | Which level of evidence settled each check |
cells/<cell>/timings.jsonl, transcripts/pass<k>.turns.jsonl | aggregate_timings.cell_rows 499 and collect 1098; aggregate_metrics.compute_cell 539 to 541 through timing_lib.pass_timing_rows | The timing columns (chapter 09) |
<task>/task_manifest.json, task_factors.json | aggregate_metrics._load_task 373 to 386; gen_cell_narrative.narrate 161 | Task type, band, obligation, edit cone, factor scores |
<run set>/run_manifest.json | gen_run_set_report.build_full_report 1677; gen_iteration_summary.ensure_article 401; cleanup_run_set._max_passes 72 | Status, validity flag, model, prompt versions, pass ceiling |
<run set>/batch_outcome.json | gen_run_set_report.build_full_report 1678; publish_experiment_results._read_outcome 190 to 201; server.py 1013 | How the batch ended; a non-terminal status refuses publication |
<run set>/closeout_progress.json | backfill_closeout._read_progress 40; server.py 1432; auto_experiment.py 253 | Which stage failed |
<run set>/results/results.md | gen_run_set_report.write_report 2108 to 2113 (the existing part two); gen_run_set_analysis.draft_run_set 568 to 570; publish_experiment_results.publish 234 to 242 (existence and freshness); server.py 1286 | The written analysis to preserve, to draft, or to publish |
<run set>/results/analysis_draft.json | gen_run_set_analysis._cells_in_earlier_draft 742 to 755; server.py 1270 | Whether an earlier draft is stale, and why a draft failed |
8. Reports/Iteration Summaries/<run>-summary.md | gen_iteration_summary.ensure_article 399, reset_draft_article 437 to 440, refresh_appendix 471; gen_run_set_analysis._draft_article 684 to 686 | Whether prose exists, whether it is a draft, where the appendix starts |
8. Reports/latest-first.md | gen_iteration_summary.update_index 483 to 484; gen_run_set_report._index_summary 2147 to 2151 | The newest-first index to add to or repair |
6. Metrics/cell_factor_matrix.csv, paired_response_matrix.csv, pass_count_matrix.csv, time_matrix.csv, token_surface_inputs.csv | aggregate_metrics.upsert_csv 819 (the existing header and rows); update_monitoring.update_monitoring 75 to 78; gen_run_set_report.attach_campaign_rows 213 to 221 and _crossrun_section 969; gen_iteration_summary._rows 63 to 79; metrics_db.import_factor_matrix 550 to 578 (the matrix only) | The cumulative campaign tables |
5. Run Sets/run_index.csv | run_batch.set_run_index_status 799 to 830; update_monitoring 79 | The campaign index; the row’s status column |
1. Harness/config/run_defaults.json | aggregate_metrics._load_defaults 107 to 117 | The budget ceiling and the pass-at-budget grid |
7. Monitoring/monitor_state.json, run_history.json, dashboard/dashboard_config.json | update_monitoring 118 to 119, 179 to 180, 215 to 216 | The existing state to preserve and extend |
7. Monitoring/fabric_events.jsonl | aggregate_timings.collect 1073 to 1075 | Model start-ups that overlap a cell’s waiting (chapter 03) |
6. Metrics/metrics.db | metrics_db.connect 364; run_set_report_metrics 790 to 800 (read-only) | The extracted behavioural tables |
| The research repository and the portal checkout | publish_experiment_results 64 to 111 through git | Whether an automatic commit and push are safe |
3.4 Network
Two stages reach outside the machine. The publication stage runs git fetch and git push against the research repository’s remote and the portal’s remote (publish_experiment_results.py 95, 106, 128, 161, 181), and the drafting stage runs the analyst model’s command-line tool, which reaches the model’s service (gen_run_set_analysis.ask_model 220 to 243, _run_subprocess 245 to 259). Nothing else in this chapter opens a connection. The publication stage also runs publish_snapshot.py as a child process to build the portal bundle (139 to 141, chapter 12).
4. The happy path in order
4.1 The batch ends and the closeout begins, run_batch.py lines 1046 to 1064 and 833 to 858 (flow-model node closeout.started)
Whatever ended the cell loop, the finally block decides the batch’s status: complete, aborted_api_errors when the circuit breaker stopped it, or stopped_before_completion with a plain-English reason from _stop_reason (1055 to 1063), and calls _closeout (1064). _closeout takes the publish lock for the whole chain (848 to 849), because serialising only the index write would leave two lanes free to overwrite each other’s aggregate and report outputs (844 to 847). _closeout_locked first writes the three cheap, certain records: active_batch.md with the final status (853; _write_active_batch 543 to 555), batch_outcome.json with the counts of cells planned, attempted, skipped and not reached and the reason (854; 617 to 650, whose docstring says why: without it a report cannot tell a run set planned as eleven cells and stopped after eleven from one that finished), and the run index row’s status (855; set_run_index_status 773 to 830, whose docstring records that 34 of the 53 rows still said active on 2026-08-31 because nothing had ever closed them). It then builds the progress record with every stage pending and writes it (856 to 857).
4.2 The stage loop, lines 859 to 895
For each stage in order: the stage is marked running with its start stamp and the record written (862 to 865); the program is invoked through the seam with its arguments (868; default_invoke 113 to 140 captures a closeout program’s output because the record keeps it as evidence); its exit status and output are read (869); the files it printed that exist under the experiment root become its artifacts (870; _printed_artifacts 751 to 770); a non-zero status marks the stage failed with the last non-empty line the program printed as its message and marks the closeout failed (871 to 875; _last_output_line 732 to 738); a zero status with a SKIPPED: line marks it skipped with that reason, and otherwise done (876 to 882); an exception in the invocation itself marks it failed with the exception’s name and text (883 to 888); and in every case the end stamp and the seconds are recorded and the record written (889 to 892). After the seventh stage the record’s status becomes failed or complete with its end stamp (893 to 895). A program that failed cannot mark itself skipped, because the status check comes first (test at test_closeout_progress.py 140).
4.3 Stage 1, aggregate_metrics.py (nodes closeout.aggregate_metrics, aggregate.process_run_set, aggregate.reached_row, aggregate.skipped, aggregate.load_cell, aggregate.reconcile_check, aggregate.compute_cell, aggregate.write_all, closeout.matrix_row)
The closeout runs the aggregator with no arguments (568), so main (1183 to 1198) lists every run set with a cells/ folder (1190 to 1192), processes each, and writes the tables once over the whole campaign. process_run_set (1098 to 1141) visits every cell folder holding a metrics record (_iter_cell_dirs 1034 to 1039). A chained step reached without an agent, recognised by its status satisfied_by_start_tree, gets its row from its own record at zero cost through reached_cell_metrics (1104 to 1111; 1046 to 1095, Decision Sheet item 54 and tracker item 2.35: the token counters are zero and the cost-to-success figures are left empty rather than zero so a per-task median never sees a free success). A cell with no token ledger is skipped with a line on standard error, because an aborted or unscored cell can leave a metrics record without one (1113 to 1120); a cell with a ledger but no scoring verdict is skipped the same way, because run set 010 aborted after its first pass spent tokens but before the scorer ran, and crashing there used to kill the whole sweep (1121 to 1130, found 2026-08-29 while backfilling 71 unaggregated cells). Every other cell is loaded (load_cell 361 to 370: the metrics record, the verdict, the ledger rows, the documentation snapshot, the task’s manifest and factors), its transcripts parsed for usage, result and compactions (1133; _pass_transcripts 223 to 234), and reconciled: reconcile_check (244 to 267) recomputes each pass’s ledger sums against the transcript’s reported usage and raises a RuntimeError naming the cell on a mismatch beyond half a percent or fifty tokens that has no remainder row, because the ledger must be regenerated and never patched (245 to 248). compute_cell (398 to 687) then produces every column of the cell’s row: the success pass from the scorer’s passes array (404), the token sums within the success scope (413 to 426), the navigation family (428 to 447), the diff statistics and outside-cone edits (447 to 449), context health and cost (452 to 472), attribution exactness and the remainder share (474 to 481), dose and cache descriptives (483 to 486), the stale-document probe (488 to 491), the visible and hidden verdicts with the hidden verdict None when the task had no scorer (493 to 499), the validity flag copied from the scorer alone (500 to 505), the timing columns and the restated new-context figures (507 to 578), shell_available (580 to 593), and the row itself (595 to 687).
write_all (1144 to 1168) upserts the matrix, adding at the tail any of the four column groups the header lacks (1152 to 1155), then the pass-count and time matrices (1156 to 1159), builds the pairings (1160; build_pairings 877 to 900 groups cells by run set, project, task, model class and seed, deliberately not by profile, because keying on the profile made cross-arm groups impossible and returned zero rows until 2026-08-06) and upserts the paired matrix and the surface inputs by the pairing key (1161 to 1164), and rewrites the two JSON files wholesale (1165 to 1168). upsert_csv (819 to 860) keeps the file’s header as the contract: a measure is never added because it turned up in the row, only when named in add_columns, which appends it at the end and pads every existing row; before 2026-09-03 a column absent from the header was silently dropped while the program reported success (831 to 835).
4.4 Stage 2, update_monitoring.py (node closeout.update_monitoring)
update_monitoring (70 to 112) reads the five tables and the run index (75 to 79), counts the completed, valid and invalid cells, the budget exhaustions and the ledger warnings (81 to 88; the warning is a proxy, the share of remainder tokens above half a percent, because the script never reads a cell’s metrics record, 24 to 31), finds the newest run set (90 to 92), and rewrites the four monitoring outputs: monitor_state.json with the existing fields preserved (115 to 135), LIVE_SUMMARY.md (138 to 173), run_history.json appended by run identifier (176 to 198), and the dashboard’s table copies and plot points (201 to 241).
4.5 Stage 3, gen_iteration_summary.py (node closeout.gen_iteration_summary)
The entry point (494 to 503) draws the run set’s figures from the campaign tables and builds the appendix through the report generator’s build_data_of_record (build_appendix 240 to 250; _draw_figures 204 to 237 writes up to three SVG files under Iteration Summaries/assets/<run>/, from gate-passing pairs only), creates the article from the skeleton only when no article exists or its writing prompts are still in place (ensure_article 397 to 421, with the design lines counted from the cell folders because the campaign table can lack a cell, 364 to 394), rewrites everything from the appendix heading to the end of the file and nothing above it (refresh_appendix 460 to 479), and adds the article to the newest-first index (482 to 491). The docstring records why the appendix is built by the report’s functions: until 2026-09-02 a second file computed the same tables from the campaign tables alone and was wrong in three ways, including reporting that no cell had ever passed when 235 of 339 had (18 to 33).
4.6 Stage 4, gen_run_set_report.py (node closeout.gen_run_set_report)
main (2195 to 2218) resolves the run set and calls write_report (2100 to 2131), which reads any existing full report (2108 to 2113), builds the new one (2117; build_full_report 1675 to 1736: the status block, the part-one marker, the provenance sentence, and eighteen numbered sections from the cells’ own records enriched by the campaign tables where rows exist, ending with the empty part two), unwraps it to one line per paragraph (2117; reflow), merges the existing part two back in (2118; _merge_preserving_part_two 2057 to 2076: part one is regenerated wholesale because it is only a rendering of the artefacts, part two is somebody’s writing and is carried forward unless it is still the untouched placeholder), writes it (2119), writes the short copy under 8. Reports/Run Set Results/ (2121 to 2127) and adds or repairs the index line (2129; 2138 to 2185). The transcript reader is called once per cell under two limits, thirty seconds per cell and 180 for the run set, because on 2026-08-31 one cell of run set 013 did not finish being read in 280 seconds and an unbounded call would put the closeout’s guarantee of finishing at the mercy of another program (1153 to 1220); cells left unread are named in the report (1206 to 1219). The two sibling hooks, extract_cell_metrics.extract_cell and metrics_db.run_set_report_metrics, are imported by name and their absence is stated in the report rather than hidden (1058 to 1073, 1400).
4.7 Stage 5, gen_run_set_analysis.py (node closeout.gen_run_set_analysis)
main (775 to 832) declines with a SKIPPED: line when drafting is switched off (796 to 799) and otherwise calls draft_run_set (547 to 672). That function reads the full report and skips, with the reason recorded, when there is none (568 to 577), when part two already holds an analysis a person wrote or accepted (580 to 587), or when it holds a draft nobody has read, unless --redraft was given or the run set has grown since the draft, in which case the stale draft is replaced and the replacement recorded (590 to 607; the growth rule exists because of run set 066 on 2026-09-02, docstring 30 to 38). It sends part one and the data of record to the model with every tool switched off (ask_model 220 to 243, two attempts and one correction when the answer lacks a required section, 623 to 639), records a failure into results/analysis_draft.json rather than only printing it, because on 2026-09-02 run set 064 finished with a complete report and an empty analysis and the only trace was a systemd journal (613 to 645), repairs em dashes and reflows the answer (647 to 649), lists every number in the draft that part one does not carry (650), splices the draft under the part-two heading with a notice naming the model and the day (652 to 659), rewrites the short copy so it stops denying an analysis exists (660 to 665; refresh_summary 2079 to 2097), fills the article’s empty prose sections without ever failing the run for it (667 to 669; _draft_article 675 to 739), and writes the provenance record (671). Deleting the notice is how a person accepts the draft; the report generator, the short copy and the interface all read that one marker (docstring 41 to 51; DRAFT_MARKER at gen_run_set_report.py 132).
4.8 Stage 6, cleanup_run_set.py (node closeout.cleanup_run_set)
main (229 to 250) calls cleanup_run_set with --apply (179 to 226). The run set must sit directly under 5. Run Sets and have a cells/ folder (51 to 60); every cell must have a scoring verdict, except one the driver deliberately left unscored because its last allowed pass died on a provider error (98 to 115; 79 to 95, the rule of 2026-09-03 that let run set 066 be cleaned, tracker item 2.26); only the eleven named directory kinds below workspace/ and staged_snapshot/ are candidates, symbolic links are never followed, and a folder named coverage is never a candidate because Project H has source under that name (29 to 41, 118 to 138). The candidates are sized, removed, and recorded in cleanup.json with the paths and bytes (184 to 225).
4.9 Stage 7, publish_experiment_results.py (node closeout.publish_experiment_results)
main (271 to 295) declines with SKIPPED: when publishing is switched off (285 to 288) and otherwise calls publish (221 to 263). The run set’s outcome must be terminal (232; _read_outcome 190 to 201), the full report must exist and must be newer than the outcome record, so a stale report is never pushed (233 to 242). The research checkout must be a repository root on the publication branch (244; 64 to 68, 90 to 94); its tip is synchronised with the remote, and only when the local commits ahead of the remote are all the publisher’s own earlier failed pushes are they pushed, otherwise the publication stops rather than pulling, rebasing or pushing unrelated commits (245 to 246; 87 to 111). The narrow set of paths (_source_paths 204 to 218: the run set, the active-batch file, the run index, the run tags, the metrics and monitoring folders, the article and its assets, the short report and the index) is added and committed with --only so another session’s staged work stays staged for its owner, then pushed (247 to 253; 114 to 129). The portal snapshot is built by publish_snapshot.py into a temporary folder (132 to 145), committed in a detached temporary worktree based on the portal’s remote tip so a designer’s uncommitted files never enter a result commit, pushed, and the worktree removed and pruned (255 to 258; 148 to 187). Both commits carry the trailer Experiment-Publisher: v1 (33), which is what lets the synchroniser recognise its own earlier commits (81 to 84).
4.10 The recovery of a failed publication after the chain, backfill_closeout.py
When the seventh stage failed, the progress record says so and the closeout’s status is failed. unresolved_run_sets (54 to 58) lists every run set whose publication stage failed; backfill (103 to 130) takes the publish lock, replays publish for each, and on success writes a recovery member into the failed stage with the two commit identifiers and the previous status and message, sets the stage done, sets the closeout complete when no stage remains failed (72 to 100), and commits the repaired progress record (119 to 125). The automatic loop keeps one warning per unresolved publication in its state file (auto_experiment.py 239 to 272, chapter 11).
4.11 The tools outside the chain
aggregate_timings.main (1137 to 1184) collects every cell’s activities, measured from timings.jsonl or reconstructed from artefacts (cell_rows 489 to 656; reconstruct 307 to 460), places the agent’s own account inside its turn (place_derived_rows 1007 to 1051), ranks the activities with the cells the scorer ruled invalid left out and counted separately (summarize_campaign 868 to 968), and writes the four timing files (1155 to 1170); the interface runs the same functions and writes the same files on its first request after a run set changes (server.py 3247 to 3278). gen_cell_narrative.main (472 to 506) writes narrative.md into one cell or every cell from the cell’s own files and changes nothing else. extract_cell_metrics.extract_cell (718 to 1031) returns a tree of behavioural records for one cell, every number a deterministic function of the preserved files, with None and a reason where a quantity cannot be determined (docstring 4 to 16); metrics_db.backfill (487 to 540) stores those records per cell with a commit per cell so an interrupted run resumes at the next cell, mirrors the matrix for round-trip checks (550 to 606), and exports eight new tables (701 to 741). reflow_markdown.reflow (22 to 143) unwraps prose while preserving fences, tables, headings, rules, front matter, callouts and hard breaks, and the three generators call it before writing (gen_run_set_report.py 62 and 2117; gen_iteration_summary.py 49 and 420; gen_run_set_analysis.py 82 and 643); lint_report_prose.main (136 to 156) scans every Markdown file under 8. Reports/ and the six write-up names beside each run set and exits 1 when three or more consecutive prose lines are found.
5. The state machines of the closeout and of the aggregator’s gate
The first diagram is the flow model’s chapter 10 subgraph: the ten closeout nodes and the eight aggregator nodes, with every edge’s line range and the record condition the model records as its witness. The chain’s between-stage edges are three per stage in the model (done, skipped, failed), drawn here once per stage with the three statuses on the label. A renderer can regenerate it from the nodes whose chapter is 10 and the edges with an endpoint among them; the one inferred witness is marked.
stateDiagram-v2 state "driver.cell_result (02)" as DCR state "closeout.started" as S0 state "closeout.aggregate_metrics" as S1 state "closeout.update_monitoring" as S2 state "closeout.gen_iteration_summary" as S3 state "closeout.gen_run_set_report" as S4 state "closeout.gen_run_set_analysis" as S5 state "closeout.cleanup_run_set" as S6 state "closeout.publish_experiment_results" as S7 state "closeout.complete" as OK state "closeout.failed" as BAD DCR --> S0 : last cell, run_batch 1043..1066 (batch_outcome.status complete; closeout exists) DCR --> S0 : unhappy, circuit breaker, 1034..1041 (status aborted_api_errors; cells_skipped) DCR --> S0 : unhappy, stop signal or crash, 1046..1066 (status stopped_before_completion) S0 --> S1 : progress written, first stage running, 857..866 (closeout_stages[aggregate_metrics].status set) state S1 { state "aggregate.process_run_set" as A0 state "aggregate.reached_row" as A1 state "aggregate.skipped" as A2 state "aggregate.load_cell" as A3 state "aggregate.reconcile_check" as A4 state "aggregate.compute_cell" as A5 state "aggregate.write_all" as A6 state "closeout.matrix_row" as A7 [*] --> A0 : aggregate_metrics.py started, 1183..1200 A0 --> A1 : reached step, 1104..1111 (metrics.status satisfied_by_start_tree; matrix row exists) A0 --> A2 : unhappy, no ledger or no verdict, 1113..1130 (no matrix row) [D3] A0 --> A3 : ordinary scored cell, 1131 (ledger, summary and metrics exist) A3 --> A4 : records loaded, 1132..1134 (stream 1 exists) A4 --> A5 : ledger reconciles, 1135 (matrix row; stage done) A5 --> A6 : every cell computed, 1136..1170 (stage done; pass_at_budget exists) A1 --> A6 : row prepared, 1108..1170 A6 --> A7 : tables upserted, 1152..1155 (matrix row exists) } S1 --> S2 : done, skipped or failed, 866..891 (closeout_stages[aggregate_metrics].status; next stage status set) A4 --> S2 : unhappy, RuntimeError mismatch beyond tolerance, aggregate_metrics 244..268 (stage failed; message matches reconcil) [inferred] A2 --> S2 : unhappy, no row for this cell, 1130 A7 --> S2 : the row exists, the chain continues, run_batch 866..891 S2 --> S3 : done, skipped or failed, 866..891 S3 --> S4 : done, skipped or failed, 866..891 S4 --> S5 : done, skipped or failed, 866..891 S5 --> S6 : done, skipped or failed, 866..891 S6 --> S7 : done, skipped or failed, 866..891 S7 --> OK : done or skipped, 866..895 (closeout.status complete) S7 --> BAD : unhappy, failed, 866..895 (closeout.status failed)
The second diagram’s states are the values of a stage’s status in the progress record and the values part_two_state returns for a report on disk, with the drafter’s transitions; both are recorded values, not the model’s nodes. Proposed identifiers for a later stage of the model: closeout.stage.pending, closeout.stage.running, closeout.stage.done, closeout.stage.skipped, closeout.stage.failed, closeout.stage.recovered, report.part_two.missing, report.part_two.empty, report.part_two.drafted, report.part_two.reviewed.
stateDiagram-v2 state "one stage of closeout_progress.json" as STAGE { [*] --> pending : _new_closeout_progress 658..682 pending --> running : started_at stamped, 862..864 running --> done : exit 0, no SKIPPED line, 882 running --> skipped : exit 0, SKIPPED line, 877..880 (message is the reason) running --> failed : exit not 0, 871..875 (message is the last printed line); or the invocation raised, 883..888 failed --> done : backfill_closeout._record_recovery 72..100 for publish_experiment_results only; recovery member keeps the old failure } state "part two of results/results.md" as P2 { [*] --> missing : no part-two heading, part_two_state 2046..2048 missing --> empty : gen_run_set_report writes the placeholder, _part_two 1647..1656, at 2119 empty --> empty : every closeout regenerates part one and leaves the placeholder, 2057..2076 empty --> drafted : gen_run_set_analysis splices a draft under the notice, 652..659 (DRAFT_MARKER 132) drafted --> drafted : a later closeout carries the draft forward, 2069..2076 drafted --> empty : reset for a redraft or growth, then redrafted, 590..607 drafted --> reviewed : a person deletes the draft notice (docstring 41..51) reviewed --> reviewed : never overwritten, 580..587 and 2057..2076 }
6. The sequence of one closeout
The participants are the batch driver, the seven stage programs, the analyst model’s tool, the two Git remotes, and the files. The interface reads the progress record while the chain runs.
sequenceDiagram participant D as run_batch._closeout_locked 852 participant L as publish_lock (publish_lock.py 38) participant X as run set and campaign files participant A as aggregate_metrics.py 1183 participant M as update_monitoring.py 244 participant I as gen_iteration_summary.py 494 participant R as gen_run_set_report.py 2195 participant N as gen_run_set_analysis.py 775 participant C as cleanup_run_set.py 229 participant P as publish_experiment_results.py 271 participant G as git remotes and the analyst model D->>L: acquire fiveb_publish.lock 848 (blocks until free) D->>X: active_batch.md 853; batch_outcome.json 854; run_index.csv status 855 D->>X: closeout_progress.json, seven stages pending 856..857 loop each stage in CLOSEOUT_CHAIN 859..892 D->>X: stage running, started_at 861..864 D->>A: default_invoke(script, args) 866 (output captured, 113..140) A->>X: read every run set's cells 1190..1196; skip loudly 1113..1130; reconcile 1134 A->>X: upsert the five tables 1152..1164; rewrite the two JSON files 1165..1168 A-->>D: exit status, stdout, stderr 867 D->>X: artifacts 868; status done, skipped or failed 869..880; finished_at, seconds 887..892 end Note over D,M: stage 2 M->>X: read the five tables and run_index.csv 75..79; write 7. Monitoring 94..103 Note over D,I: stage 3 I->>X: figures from the campaign tables 204..237; article skeleton if none 397..421; Appendix A rewritten 460..479; latest-first.md 482..491 Note over D,R: stage 4 R->>X: read cells, manifest, outcome, campaign rows 1675..1680; extractor hooks under 180 s 1153..1220 R->>X: results/results.md with part two preserved 2117..2119; Run Set Results short copy 2121..2127; index line 2129 Note over D,N: stage 5 N->>X: read results.md 568..570; decide by part_two_state 580..607 N->>G: claude -p --tools "" 230..232, up to two attempts plus one correction 623..639 G-->>N: part two text N->>X: results.md part two with the draft notice 652..659; short copy 660..665; article prose 667..669; analysis_draft.json 671 Note over D,C: stage 6 C->>X: refuse unless every cell is scored 98..115; remove the named folders 190..212; cleanup.json 225 Note over D,P: stage 7 P->>X: batch_outcome.json terminal 232; results.md newer than the outcome 233..242 P->>G: git fetch, push own earlier commits only 87..111; commit --only the result paths, push 114..129 P->>X: publish_snapshot.py into a temporary folder 132..145 P->>G: detached worktree on the portal's remote tip, commit, push 148..187 P-->>D: exit status and the JSON result 294 D->>X: closeout status complete or failed, finished_at 893..895 D->>L: release the lock
7. The records, with their writers and readers
Every file the chain and the four tools write. Cell-level records are chapters 06, 08 and 09’s and are cited there. Readers are named by file and line where the reading is one call and by file where it is spread. Every record here is single agent only.
| Record | Writer | Fields or content | Readers |
|---|---|---|---|
<run set>/closeout_progress.json (flow-model key closeout; its stages list is closeout_stages) | run_batch._write_closeout_progress 685 to 719, from _new_closeout_progress 658 to 682 and the loop 861 to 895; backfill_closeout._record_recovery 72 to 100 adds recovery and recovered_at | schema_version (5b.closeout_progress.v001), run_id, batch_status, status (running, complete, failed), started_at, finished_at, stages (each with key, title, what, script, status in pending, running, done, skipped, failed, started_at, finished_at, seconds, message, artifacts, and after a recovery recovery with recovered_at, method, source_commit, website_commit, previous_status, previous_message) | backfill_closeout._read_progress 40 and unresolved_run_sets 54; server.py _closeout_progress 1432 to 1480 (validated field by field; a damaged record reads as none) into the live view at 2176; auto_experiment.py 253 to 271 |
<run set>/batch_outcome.json (batch_outcome) | run_batch._write_batch_outcome 617 to 650; also attest_batch_outcome.py 145 to 149 by hand, with record_provenance operator_attested | schema_version (5b.batch_outcome.v002), record_provenance, run_id, status (complete, aborted_api_errors, stopped_before_completion), stopped_before_completion, reason, cells_planned, cells_attempted, cells_skipped, cells_not_reached, recorded_at | gen_run_set_report.build_full_report 1678 and build_summary 1743; publish_experiment_results._read_outcome 190 to 201; server.py _batch_outcome 1013 to 1020 |
5. Run Sets/active_batch.md (active_batch) | run_batch._write_active_batch 543 to 555, at start and at closeout 853 | Run set, Cells, Status, and Started, Stopped or Completed with the stamp | backfill_turn_ledgers.batch_is_running 59 (chapter 09); server.py _parse_active 736; publish_experiment_results._source_paths 207 (committed) |
5. Run Sets/run_index.csv, the status column (run_index; chapter 06 writes the row) | run_batch.set_run_index_status 773 to 830, under the lock | The batch’s final status in the matching row; every other column and row written back as read | update_monitoring 79; server.py; publish_snapshot.py |
6. Metrics/cell_factor_matrix.csv (matrix) | aggregate_metrics.write_all 1152 to 1155 through upsert_csv 819 to 860, keyed by run_id and cell_id | Every column compute_cell 595 to 687 or reached_cell_metrics 1046 to 1095 produces, in the header’s order: identity, model provenance, the prompt and feedback factors, the handoff columns, the factor scores, the gates and the primary metric, the token splits, the navigation family, the diff statistics, context health, cost, the timing columns and the restated new-context columns (metric_definitions.md section 7) | update_monitoring 75; gen_run_set_report.attach_campaign_rows 213 and _crossrun_section 969; gen_iteration_summary._rows 216; metrics_db.import_factor_matrix 550; analyze_coverage.load_rows and server.py (chapter 11); publish_snapshot.py and paired_analysis.py (chapter 12) |
6. Metrics/pass_count_matrix.csv, time_matrix.csv (pass_count_matrix, time_matrix) | write_all 1156 to 1159, same key | The header’s subset of the same row: passes and budget exhaustion; seconds to success and lease wait | update_monitoring 77; gen_run_set_report.attach_campaign_rows 220 to 221 |
6. Metrics/paired_response_matrix.csv, token_surface_inputs.csv (paired_response_matrix, token_surface_inputs) | write_all 1160 to 1164 through build_pairings 877 to 900 and _pairing_row 903 to 946, keyed by run_id, task_id, lap_profile_id, model_class, seed | One row per documented arm against its best co-located baseline: the deltas of tokens, seconds and passes, the docs-subtracted delta, the per-category deltas, the validity flag; the surface row projects token_delta with empty axes (1171 to 1180) | update_monitoring 76 and 78; gen_run_set_report.attach_campaign_rows 219; gen_iteration_summary._draw_figures 217; paired_analysis.py (chapter 12) |
6. Metrics/pass_at_budget.json, group_analysis.json (pass_at_budget, group_analysis) | write_all 1165 to 1168, rewritten wholesale from pass_at_budget 955 and group_analysis 975 | Pass rates at each budget of the grid; the per-group success and dispersion under the ceiling; both over the last aggregation’s cells only | gen_run_set_report.py |
7. Monitoring/LIVE_SUMMARY.md, monitor_state.json, run_history.json, dashboard/tables/*.csv, dashboard/dashboard_config.json, plots/plot_data.json | update_monitoring 138 to 173, 115 to 135, 176 to 198, 208 to 211, 223 to 224, 241 | The headline counts and the newest-first run list; the state with active_batch preserved; one entry per run set keyed by identifier; byte copies of the five tables; the plot points with empty axes | the operator; server.py; publish_experiment_results._source_paths 211 (committed) |
8. Reports/Iteration Summaries/<run>-summary.md and assets/<run>/fig1-paired-deltas.svg, fig2-category-deltas.svg, fig3-dose-consumed.svg | gen_iteration_summary.ensure_article 420 (the skeleton, once), refresh_appendix 477 to 478 (from the appendix heading down, every closeout), reset_draft_article 441 to 448 (a redraft); the figures at 228 to 235; the prose sections by gen_run_set_analysis._draft_article 731 | The article’s prose halves with the draft notice or a person’s writing, and ## Appendix A. Data of record built by gen_run_set_report.build_data_of_record 1940 to 2008 | gen_run_set_analysis._draft_article 684; server.py (the reports view); publish_experiment_results._source_paths 212 to 214 |
8. Reports/latest-first.md | gen_iteration_summary.update_index 482 to 491; gen_run_set_report._index_summary 2138 to 2185 | One line per article and per short report, newest first, with the original date kept | readers of the reports folder; committed at 217 |
<run set>/results/results.md and results/figures/*.svg | gen_run_set_report.write_report 2119 (part one regenerated, part two preserved), the figures by _figures_section 929; part two by gen_run_set_analysis.draft_run_set 658 | The status block, part one’s eighteen numbered sections each naming its source, and part two’s five sections (138 to 155) as a placeholder, a draft under the notice, or a person’s analysis | gen_run_set_analysis.draft_run_set 568; publish_experiment_results.publish 234 to 242; server.py _full_report_analysis 1278 to 1300; lint_report_prose.report_paths 121 to 123 |
8. Reports/Run Set Results/<run>-results.md | gen_run_set_report.write_report 2121 to 2127 and refresh_summary 2079 to 2097 | The short copy, with the status callout and the outcome lines | server.py; committed at 215 to 216 |
<run set>/results/analysis_draft.json | gen_run_set_analysis._write_record 758 to 762, at 644 and 671 | run_id, model, drafted_at, script, wrote, skipped, em_dashes_repaired, numbers_not_in_part_one, cells_drafted, and when applicable declined_reason, redrafted_because, failed, executable, draft, article_reset, article_sections_written | _cells_in_earlier_draft 742 to 755; server.py 1259 to 1275 (why a draft failed) |
<run set>/cleanup.json | cleanup_run_set._write_cleanup_record 155 to 176, at 209 and 225 | schema_version (5b.cleanup.v001), run_id, applied, directories_removed, bytes_freed, removed_paths, directories_found, bytes_found, candidate_paths, and on a failed removal failed | the operator; committed with the run set |
The research repository commit Publish experiment results for <run> and the portal commit Publish experiment website update for <run> | publish_experiment_results._commit_paths 114 to 129 and _publish_portal_snapshot 148 to 187; the repaired progress record by backfill_closeout.backfill 120 to 125 | The paths of _source_paths 204 to 218; the folder experiment-interface/ of the portal; both with the trailer Experiment-Publisher: v1 | _automatic_commits_only 81 to 84 on the next publication; the portal’s build and deployment (chapter 12) |
6. Metrics/cell_timings.csv, stage_timings.csv, pass_timing.csv, timing_summary.json | aggregate_timings.main 1155 to 1170; server.py _timing_data 3272 to 3278 | One row per cell with the activities in named columns; one row per activity with source measured or reconstructed; one row per attempt (timing_lib.PASS_TIMING_COLUMNS); the campaign ranking with the excluded cells counted separately (943 to 968) | server.py 3280 to 3363; the B4 forensics |
cells/<cell>/narrative.md | gen_cell_narrative.main 500 to 501 | A chronological account of the cell from its own files, with the retry notices as the runner’s constants (35 to 45) | the operator |
6. Metrics/metrics.db and the eight exported tables nav_cell_metrics.csv, nav_pass_metrics.csv, doc_artifact_reads.csv, doc_artifact_frequency.csv, doc_artifact_frequency_by_condition.csv, pass_shape.csv, condition_compare.csv, model_compare.csv | metrics_db.backfill 487 to 540 (per cell, store_cell 433 and log_skip 453), import_factor_matrix 550 to 578, export_csvs 724 to 741 | The tables of the schema at 59 to 356 (cells, passes, file_access, file_edits, shell_commands, searches, tool_call_counts, doc_artifact_access, file_composition, runsets, extraction_log, schema_meta, the mirrored matrix, and eight views) | metrics_db.run_set_report_metrics 774 to 879 into the report’s extended section (gen_run_set_report.py 1400); crosscheck_artifact_use 613 |
The lineage from the cell records through the chain to the published portal is drawn below. Every arrow is a writer or reader relationship from the table above or from section 3.3.
flowchart LR subgraph cell [cells/<cell>/ (chapters 06, 08, 09)] M[metrics.json] S[scoring_summary.json] LG[token_ledger.jsonl] TR[transcripts/pass k .stream.jsonl] P[patches/pass k .patch] SN[lap_path_manifest_snapshot.json] CM[cell_manifest.json] HT[hidden_tier_result.json] TI[timings.jsonl, pass k .turns.jsonl] end subgraph driver [run_batch.py, under publish_lock 848] AB[active_batch.md 553] BO[batch_outcome.json 645] RI[run_index.csv status 773..830] CP[closeout_progress.json 685..719] end AG[aggregate_metrics.py<br/>process_run_set 1098; compute_cell 398; write_all 1144] M --> AG S --> AG LG --> AG TR --> AG P --> AG SN --> AG CM --> AG TI --> AG CFM[(6. Metrics/cell_factor_matrix.csv 1152)] PCM[(pass_count_matrix.csv 1156)] TM[(time_matrix.csv 1158)] PRM[(paired_response_matrix.csv 1163)] TSI[(token_surface_inputs.csv 1164)] PAB[(pass_at_budget.json 1165, group_analysis.json 1167)] AG --> CFM AG --> PCM AG --> TM AG --> PRM AG --> TSI AG --> PAB UM[update_monitoring.py 70] CFM --> UM PRM --> UM PCM --> UM TSI --> UM RI --> UM MON[(7. Monitoring/LIVE_SUMMARY.md, monitor_state.json,<br/>run_history.json, dashboard/tables, plots/plot_data.json)] UM --> MON IS[gen_iteration_summary.py 494<br/>figures 204; appendix via build_data_of_record 1940] CFM --> IS PRM --> IS M --> IS S --> IS ART[(8. Reports/Iteration Summaries/run-summary.md, assets/run/fig*.svg)] IS --> ART IDX[(8. Reports/latest-first.md)] IS --> IDX RP[gen_run_set_report.py 2195<br/>build_full_report 1675; merge part two 2057] M --> RP S --> RP CM --> RP HT --> RP BO --> RP CFM --> RP PRM --> RP TR -->|extract_cell hook 1153| RP DB[(6. Metrics/metrics.db<br/>metrics_db.backfill 487)] DB -->|run_set_report_metrics 774| RP RES[(run set/results/results.md, figures/)] SHORT[(8. Reports/Run Set Results/run-results.md)] RP --> RES RP --> SHORT RP --> IDX AN[gen_run_set_analysis.py 775<br/>ask_model 220] RES --> AN ART --> AN AN --> RES AN --> SHORT AN --> ART AD[(results/analysis_draft.json 758)] AN --> AD CL[cleanup_run_set.py 229] S --> CL M --> CL CJ[(cleanup.json 155)] CL --> CJ PB[publish_experiment_results.py 271<br/>publish 221] BO --> PB RES --> PB GITR[(research repository commit 114..129)] GITP[(portal commit experiment-interface/ 148..187)] PB --> GITR PB --> GITP PS[publish_snapshot.py, chapter 12] CFM --> PS PS --> GITP BF[backfill_closeout.py 103] CP --> BF BF --> GITR BF --> CP AT[aggregate_timings.py 1137, run by server.py 3253] TI --> AT CM --> AT CT[(6. Metrics/cell_timings.csv, stage_timings.csv,<br/>pass_timing.csv, timing_summary.json)] AT --> CT SV[server.py: _closeout_progress 1432; _batch_outcome 1013;<br/>_full_report_analysis 1278; timing views 3280] CP --> SV BO --> SV RES --> SV CFM --> SV CT --> SV
8. The loops and the waits
The stage loop (run_batch.py 859 to 892) runs exactly seven iterations and never retries a stage; each invocation blocks until the program exits, with no timeout, because the closeout must finish and a stage that hangs is a stage that must be found and fixed, not skipped. Inside the stages: the aggregator’s loop is over every run set and every cell (1190 to 1196, 1102), unbounded by design and reading every transcript of every scored cell on every closeout, which is the 44.9 seconds the stage took on run set 211 (its progress record); the report’s transcript reader is bounded to thirty seconds per cell and 180 per run set (1153 to 1220) and names the cells it left unread; the drafter makes at most three model calls (two attempts, 220 to 243, plus one correction, 624 to 631), each bounded by the timeout of 900 seconds (114), which is the 132.3 seconds the stage took on run set 211; the publication’s Git calls have no timeout and the snapshot build has none (139 to 141); the publish lock blocks without limit until the other lane’s chain finishes (publish_lock.py 50 to 55). The monitoring update, the article and the cleanup loop over files only. The database backfill commits per cell so an interrupted run resumes at the next cell (512 to 516) and skips a cell already extracted by the same extractor version when asked (504 to 507). The timing aggregator’s collect walks every run set and cell (1077 to 1113) and the interface caches its result by the newest run set’s modification time so a finished batch invalidates it on its own (server.py 3210 to 3226).
9. The guards and refusals
| Guard | Where | What it refuses or records |
|---|---|---|
| The closeout runs however the batch ended | run_batch.py 1046 to 1064 | A stop or a crash still writes the outcome, the index row and every stage’s record |
| One lock around the whole chain | 833 to 849; publish_lock.py 38 to 55 | Two lanes never interleave their writes into the campaign’s files |
| A failed stage does not stop the chain | 871 to 875, 883 to 888 | The stage is failed, the closeout is failed, the next stage runs |
| A program may decline, never fail silently | 877 to 880 and SKIPPED_PREFIX 598 | Only an exit of zero with a SKIPPED: line is skipped; a failed program cannot mark itself skipped |
| The progress record is replaced atomically and never blocks the chain | 685 to 719 | A temporary file moved into place; a disk failure is printed and ignored |
| The index row is closed under the lock | set_run_index_status 773 to 830 | Only the matching row’s status column changes; a concurrent lane’s change survives |
| A cell without a ledger or a verdict gets no row, loudly | aggregate_metrics.process_run_set 1113 to 1130 | The skip is printed on standard error naming the cell |
| A reached step is a row at zero cost, never a free success | reached_cell_metrics 1046 to 1095 | Cost-to-success figures empty, counters zero, agent_ran false |
| A ledger that does not reconcile stops the aggregator | reconcile_check 244 to 267 | RuntimeError naming the cell and pass; the ledger must be regenerated, never patched |
| A hidden verdict is never invented | compute_cell 493 to 499 | hidden_pass is None when the task had no scorer |
| Validity is the scorer’s verdict alone | 500 to 505 | Since 2026-09-03 the aggregator adds no rule of its own |
| The header is the contract; a column is added only by name | upsert_csv 819 to 860 | A measure absent from the header is dropped unless named in add_columns, which appends it at the end |
| A pairing never mixes profiles wrongly | build_pairings 877 to 900 | The group key excludes the profile identifier; each documented arm pairs against every co-located baseline |
| A campaign table that does not exist is not an error | gen_iteration_summary._rows 63 to 79; gen_run_set_report._csv_rows 192 to 200 | A missing table yields no rows, and the document is still written |
| The article’s prose is never overwritten | ensure_article 399 to 400; refresh_appendix 460 to 479; reset_draft_article 437 to 440 | Only the appendix is rewritten; a draft is reset only while it carries the notice |
| Part two of the report is never overwritten | _merge_preserving_part_two 2057 to 2076; _find_part_two 2020 to 2032 searches the legacy headings too | Anything written under either heading survives regeneration; thirty-seven reports carry the old wording (96 to 103) |
| The transcript reader is bounded | collect_cell_extras 1153 to 1220; _cell_deadline 1102 | Thirty seconds per cell, 180 per run set; unread cells are named |
| A sibling hook that raises is reported, not propagated | 1194 to 1203 | The cell contributes no extended numbers and the report says so |
| The drafter never touches part one | draft_run_set 609 to 611, splice_part_two 395 | Only the text under the part-two heading is replaced |
| An analysis a person stands behind is never overwritten | 580 to 587 | reviewed skips with the reason |
| A draft nobody has read is replaced only on request or on growth | 590 to 607 | --redraft, or more cells now than the draft described |
| The model sees the prompt and nothing else | ask_model 230 to 232 | Every tool off, safe mode, a fixed system prompt |
| A drafting failure is written into the run set | 640 to 645 | analysis_draft.json carries failed; exit 2 |
| An answer missing a section is corrected once, then refused | 626 to 639 | RuntimeError after the second failure |
| Numbers the report does not carry are listed | unverified_numbers 284 to 308, at 645 | The reviewer is told what to check first |
| The article draft never costs the analysis | _draft_article 675 to 683 | A failure there is recorded as skipped and the run succeeds |
| Cleanup touches only named folders under two parents | cleanup_run_set.py 29 to 43, 118 to 138 | Never coverage, never a symbolic link, never outside workspace/ and staged_snapshot/ |
| Cleanup refuses an unscored cell | _require_scored_cells 98 to 115 | Unless the cell spent every pass on a provider error (79 to 95) |
Cleanup refuses a run set outside 5. Run Sets | _validated_run_set 51 to 60 | The parent folder must be named exactly that |
| Publication requires a terminal outcome and a fresh report | publish 232 to 242 | A non-terminal status, a missing report, or a report older than the outcome refuses |
| Publication never pulls, rebases or pushes unrelated work | _synchronise_tip 87 to 111; _require_empty_index 71 to 78; _commit_paths 126 | Only the publisher’s own earlier commits are pushed ahead; only the result paths are committed |
| The portal commit never carries a designer’s files | _publish_portal_snapshot 148 to 187 | A detached worktree on the remote tip, removed and pruned afterwards |
| Publication can be switched off | _enabled 266 to 268 | SKIPPED: with the setting named |
| A recovery keeps the old failure | backfill_closeout._record_recovery 72 to 100 | The recovery member records the previous status and message |
| The interface trusts only a well-formed record | server.py _closeout_progress 1441 to 1480 | A damaged record reads as no record |
| The timing ranking leaves out invalid cells and says so | aggregate_timings.summarize_campaign 880 to 885, 949 to 953 | excluded_cells and excluded_seconds |
| A measured cell is never bounded by a reconstruction cap | aggregate_timings.cell_rows 499 to 509 | Caps apply to reconstructed spans only (test at test_aggregate_timings.py 197) |
| The database never damages the matrix | regenerate_factor_matrix 581 to 606 | Written only when byte-identical |
| Generated prose is one line per paragraph | reflow at gen_run_set_report.py 2117, gen_iteration_summary.py 420 and 478, gen_run_set_analysis.py 643 | lint_report_prose.py fails a report that wraps |
10. Every unhappy path, in four parts
Each row states what the step is supposed to do and why it works that way, what goes wrong with the trigger and the code path, what is written with the status and exit, and what it costs downstream.
| Supposed to do, and why it works that way | What goes wrong: trigger and code path | What is written; status and exit | What it costs downstream |
|---|---|---|---|
| Run every stage and record each, so a batch that was stopped still gets its tables and reports | A stage’s program exits non-zero or its invocation raises; run_batch.py 871 to 875, 883 to 888 | The stage failed with the program’s last printed line or the exception; the closeout’s status failed; the next stage runs | The stages after it run against whatever the failed one left; the interface shows the failed stage and its message (server.py 2176); a failed publication is replayable, a failed aggregation is not and must be re-run by hand with the full sweep (6. Metrics/README.md) |
| Keep the progress record current while the chain runs | The record cannot be written; _write_closeout_progress 710 to 719 | A line on standard error; the chain continues; the record on disk is the last one written | The interface shows a stale stage; test at test_closeout_progress.py 193 |
| Close the run index row from the batch’s final status | The index is missing, unreadable, has no run_id or status column, or cannot be written; set_run_index_status 799 to 830 | Nothing; returns false | The row keeps saying active, the state 34 of 53 rows were in on 2026-08-31 (781 to 785) |
| Give every scored cell a row in the campaign tables | The cell has no ledger or no verdict; process_run_set 1113 to 1130 | A line on standard error; no row | The cell is absent from every table and from the Coverage Map (B2-count-provenance.md section 1); the report still counts it from its own files (gen_run_set_report.py 34 to 36) |
| Reconcile every pass’s ledger against its transcript before the row is written | A field’s sum differs beyond tolerance with no remainder row, including the case chapter 09 section 10 describes where the runner’s and the aggregator’s readings of the reported usage differ; reconcile_check 262 to 267 | RuntimeError; the aggregator exits with a traceback and writes no table | The whole first stage fails for every run set in the sweep, not only the offending cell; the flow model marks this edge inferred because the stage’s message text was not exercised on a fixture (open item 12). No campaign closeout has been seen to fail this way |
| Refresh the monitoring files from the tables | A table is missing; _read_csv 46 to 50 | Empty lists; the counts are zero; the files are written | LIVE_SUMMARY.md reports zero cells until the next aggregation; the README’s note that the two counts will not match records the same class of gap |
| Write the article once and rewrite only its appendix | The article exists with prose, or with a draft; ensure_article 399 to 400 | The prose is left; the appendix from its heading down is replaced (460 to 479) | Nothing lost; an article from before 2026-09-02 has its embed of the retired data file replaced and its pointer repaired (466 to 475) |
| Write part one of the report and preserve part two | The existing report carries a legacy heading, or the placeholder in its old bolded form; _find_part_two 2020 to 2032, _merge_preserving_part_two 2069 to 2074 | The legacy heading is recognised and the writing carried; a placeholder written either way is treated as empty | Nothing lost; deleting the legacy constants would make the generator throw thirty-seven reports’ analyses away (96 to 103) |
| Read every cell’s transcript for the report’s extended measures | A cell takes longer than thirty seconds, the run set longer than 180, or the reader raises; collect_cell_extras 1188 to 1219 | The cell contributes no extended numbers and the report names it; exit 0 | The extended section is incomplete for that run set; the variables to raise the limits are named in the report |
| Draft the analysis and mark it as a draft | The model’s tool is not found, exceeds the timeout, exits non-zero, returns nothing twice, or returns an answer missing a section after one correction; ask_model 220 to 243, _run_subprocess 245 to 259, 623 to 639 | analysis_draft.json with failed and the executable; a line on standard error; exit 2; the stage failed | The report keeps its placeholder; the interface says the analysis was attempted and could not be written, which is a different fact from nobody having written one (server.py 1259 to 1275); run set 064 on 2026-09-02 is the incident (the comment at 613 to 622) |
| Never overwrite a judgment | Part two is reviewed, or drafted without --redraft and without growth; 580 to 607 | SKIPPED: with the reason; exit 0 | Nothing; the draft or the analysis stands |
| Delete only regenerable folders, only after every cell is scored | A cell has no verdict and did not spend its passes on a provider error; a folder cannot be removed; _require_scored_cells 110 to 114; 196 to 210 | CleanupRefused; cleanup.json with failed when a removal failed part way; exit 2; the stage failed | The run set keeps its dependency trees (hundreds of megabytes per cell); before 2026-09-03 run set 066 could never be cleaned for this reason (86 to 88) |
| Commit the results and refresh the portal, only when it is safe | The outcome is not terminal, the report is missing or older than the outcome, a checkout is not a repository root or is on another branch, its index holds staged changes, its tip diverged with commits that are not the publisher’s, none of the expected files exist, the snapshot builder fails, or a Git command fails; publish 227 to 258 and the helpers 45 to 187 | PublishError printed; exit 1; the stage failed with that line as its message; the research commit may already have been pushed when the portal step fails (247 to 258) | The results stay unpublished until backfill_closeout.py replays the stage; run set 211’s record shows the real case, a failure at git worktree remove --force after the research commit was made, which the recovery records and repairs (72 to 100) |
| Replay a failed publication | The replay fails again; backfill 127 to 128 | The failure listed in the summary; exit 1; the progress record unchanged | The warning stays in the automatic loop’s state (auto_experiment.py 239 to 272) |
| Keep the two JSON files consistent with the CSV tables | The aggregator is run for one run set with --run-set; write_all 1165 to 1168 | pass_at_budget.json and group_analysis.json describe that run set’s cells only | A reader of the two files sees the last aggregation’s cells, not the campaign (6. Metrics/README.md); the closeout runs with no argument, so this happens only by hand |
| Count the ledger warnings without reading a cell’s record | The proxy differs from the runner’s flag; update_monitoring.py 24 to 31 | ledger_warning_count from the remainder share | The count is not the number of cells the runner flagged, and the monitoring README says most warnings attach to cells before the parser correction of 2026-08-28 |
| Rank the campaign’s time without invalid cells | A row has no validity flag; _counts_towards_results 853 to 865 | The row counts | Only a hand-built fixture lacks the flag; the harness always writes it |
| Rebuild a cell’s timings from its artefacts when it has no timing file | A reconstructed span exceeds its plausibility cap, has stamps that run backwards, or an unknown end; reconstruct 307 to 460 and RECONSTRUCTION_CAP_SECONDS 96 | The span is dropped and counted in discarded_implausible (1166 to 1170) | A re-score weeks later would otherwise read as a ten-day cell (test at test_aggregate_timings.py 376) |
| Extract behavioural measures from a legacy run set | A cell crashes the extractor; metrics_db.backfill 518 to 527 | extraction_log row with the reason; the cell skipped; the run set’s counts recorded | A legacy cell (run sets before 031) is skipped without a traceback; a current one carries its traceback |
11. The metrics this chapter produces and where each goes
This chapter is where every campaign number is written. cell_factor_matrix.csv carries one row per admitted cell with the columns of metric_definitions.md section 7: the gate (visible_pass, hidden_pass, valid_for_primary_analysis), the primary metric tokens_to_success, the token splits and the docs-subtracted variant (section 3), artifact_use_score (section 4), the pre-registered additions of section 9 (regression, cone navigation, navigation efficiency, outside-cone edits, context health, attribution exactness, dose, cost, the stale-document probe), the timing columns from chapter 09, and the four factor columns of 2026-09-01 and 2026-09-02. paired_response_matrix.csv carries the deltas of each documented arm against its best co-located baseline, still by the retired best-baseline selection that the definitions’ migration note of 2026-08-11 says the script must move off (_pairing_row 903 to 946). pass_count_matrix.csv and time_matrix.csv carry the pass and time subsets. pass_at_budget.json carries the pass rate at each budget of the grid and group_analysis.json the per-group success and dispersion under the ceiling. The monitoring files carry the headline counts. The report prints every cell’s disposition, correctness, invalidity reasons, hidden-check levels, passes, category shares, timing, resources, paired differences, cross-run context, deviations and provenance from the cells’ own files, and the article’s appendix carries the same tables. The timing tables of aggregate_timings.py carry the campaign’s ranking of where its hours went. The database’s exports carry the per-artefact read counts and the per-pass navigation measures that the report’s extended section prints when the database exists. paired_analysis.py, analyze_coverage.py, publish_snapshot.py and the interface (chapters 11 and 12) read the matrix and the paired matrix from here and nowhere else.
12. The tests that exercise the mechanism
test_closeout_progress.py (seven tests, lines 28 to 193) covers the record written before any stage, a stage moving through running and done, a failed stage keeping its message while later stages run, a declining program recorded as skipped, a failed program unable to mark itself skipped, a stopped batch still writing the record, and an unwritable record not stopping the chain. test_run_set_report.py (fifty-four tests, 143 to 709) covers the closeout under a stop, the stop reason, the circuit breaker’s outcome, a failing step not stopping the others, the appendix counting a cell the table never recorded, the appendix rewrite never touching prose, the migration of a two-file article, the two ruled locations, one physical line per paragraph, the index written once, part one naming its sources, part two present and empty, a written analysis surviving regeneration, a report complete without the campaign tables, an empty run set, and the sibling hooks named, used, tolerated and timed. test_run_set_analysis.py (thirty-one tests, 100 to 560) covers the draft and its marker, part one untouched, every tool switched off, the short copy, the provenance record, a reviewed analysis never overwritten, a draft not silently replaced, the redraft and the growth rule, the article filled and reset, a missing report, a missing section retried then refused, em dashes repaired, unverified numbers flagged, the switch, the tool search, and a failure written into the run set. test_aggregate.py (thirty tests, 131 to 805) covers the token splits, the documentation-use score, the navigation family, the reconciliation abort (359), the best-baseline deltas, pass at budget, idempotent writes (441), the full run set end to end, a ledgerless cell skipped (486), the timing columns, the new-context columns and the header extension. test_update_monitoring.py (four tests, 74 to 124), test_cleanup_run_set.py (nine tests, 35 to 168, including the provider-error exception at 94 and 110), test_publish_experiment_results.py (six tests, 84 to 219, including the publication with unrelated staged work left alone, the detached portal worktree, the backfill scanner, the required report and the switch), test_reflow_markdown.py (four tests) and test_lint_report_prose.py (three tests) cover the remaining stages and tools. test_run_batch.py 62 chains cells and runs the whole closeout once; test_lane_coordinator.py 337 checks that a concurrent closeout preserves both lanes’ rows; test_stop_and_evidence.py 164 checks the aggregator’s skip of an unscored cell; test_aggregate_timings.py (lines 110 to 604) covers the reconstruction, the caps, the ranking and the exclusion of invalid cells; the interface’s test_auto_experiment.py 1428 covers the backfill ledger.
Against section 10: every row except three is covered by a named test above. No test covers the aggregator’s reconciliation message reaching the progress record (the inferred witness), the publication’s failure between the research push and the portal push (run set 211’s case), or the monitoring proxy’s disagreement with the runner’s flag.
13. The dated incidents that shaped the code
The chain moved into the finally block on 2026-08-31 after run set 041’s eleven verdicts were lost to a deliberate stop and run set 038’s stop left active_batch.md saying running for a batch with no process behind it (run_batch.py 17 to 24, 1049 to 1054). The run index closer dates from the same day, when 34 of 53 rows still read active (781 to 785). The two report locations and the two-part rule were ruled on 2026-08-31 (gen_run_set_report.py 8 to 9), the em-dash markers were replaced the same day with the legacy wordings kept because thirty-seven reports carry them (96 to 108), the index repair dates from then (2141 to 2145), and the transcript reader’s limits from run set 013’s 280-second cell (1162 to 1173). The analysis drafter dates from 2026-09-01, when thirty-nine run sets carried an empty part two and fifteen in a row had no recorded judgment (gen_run_set_analysis.py 19 to 22); its growth rule from run set 066 on 2026-09-02 (30 to 38); its failure record from run set 064 the same day (613 to 622). The article’s measured half was merged into the article on 2026-09-02 after thirty-three data files had reported that no cell had ever passed when 235 of 339 had (gen_iteration_summary.py 18 to 33; gen_run_set_report.py 1944 to 1956). The aggregator’s loud skips date from run set 003’s wiring-gap abort and from run set 010, found on 2026-08-29 while backfilling 71 unaggregated cells (1114 to 1129); its pairing key from run set 006 on 2026-08-06, the first run set to co-locate both arms (877 to 886); its header extension from 2026-09-03 (831 to 838); its reached-step row from tracker item 2.35 on 2026-09-03 (1046 to 1055); its timing and new-context columns from 2026-09-03 (507, 542); and its validity rule from Decision Sheet item 51 on 2026-09-03 (500 to 505). The cleanup’s provider-error exception dates from 2026-09-03 and tracker item 2.26, when run set 066 could not be cleaned (cleanup_run_set.py 82 to 88). The publish lock and the whole-chain locking are the lane coordinator’s LP-16 (publish_lock.py 1 to 22; run_batch.py 786 to 793, 844 to 847). The narrative tool’s retry notices were v003 text until 2026-08-22, which reintroduced in the narrative the leak amendment A4 had removed from the notice (gen_cell_narrative.py 27 to 34). The metric definitions’ migration note of 2026-08-11 records that the paired delta still uses the retired best-baseline selection under the statistical plan’s version 5.
14. The weakest claim, what was not checked, and the token line
The weakest claim
The weakest claim is the reader list of the campaign tables in section 7 for the interface and the analysis scripts: chapters 11 and 12 have not been written, server.py changed since the flow model’s commit, and the readers named there (analyze_coverage.load_rows, publish_snapshot.py, paired_analysis.py) are taken from the results-integrity memo B2-count-provenance.md and the flow model’s record entries rather than re-read line by line in this pass. The second weakest is the statement in section 10 that no campaign closeout has failed on the reconciliation check; it rests on the flow model’s note that the message was not exercised on a fixture and on the one progress record read (run set 211), not on a scan of every run set’s record.
14.3 What was not checked
Whether every run set’s closeout_progress.json on disk validates under the interface’s field-by-field check (server.py 1441 to 1480); one record was read. Whether the eighteen section builders of the report (308 to 1645) read only the files section 3.3 names; they were read by signature, not whole. Whether aggregate_timings.reconstruct (307 to 460) and the extractor’s extract_cell (718 to 1031) have exits beyond the ones their docstrings and tests name; both were read by region. The eight database views (265 to 356) were not read. metric_definitions.md was read for section 7 only.
14.4 Token line
The session that wrote this chapter had consumed 817,201 tokens of context by the time writing began, measured as the difference of the remaining-token counter (15,000,000 at the start of the session, 14,182,799 at the last reading before the file was written); of that, 628,458 had been spent before chapter 01 was written, and the remainder on this chapter’s materials: the flow model’s chapter 10 subgraph, the draft’s outline and sampled sections, the batch driver’s closeout region in full, the seven stage programs’ entry points and main functions, the four tools’ entry points, the metric definitions’ table schema, the two folder READMEs, one real progress record, and the tests. No harness script was run and no model tokens were spent by the harness.