1. The purpose and the position in the life of the experiment
A published number is a number that can be followed from a recorded cell, through an analysis rule, to a named claim and then to a read-only page. This chapter describes that path. A cell is one recorded execution of one task, condition, phase and seed. A condition is one version of the product presented to the agent. A seed is the repetition label used to match otherwise corresponding cells. A run set is a batch of such cells stored together. The analysis path reads those records after execution and scoring. It does not run an agent, repair a product tree, or decide that an unregistered proposition is a result.
The first program, 5. Experiment/1. Harness/scripts/analyze_coverage.py, is a coverage reader. Coverage means how many valid cell rows exist for each build and task combination. The program can print grids, list gaps below a repetition target, or write a draft batch specification for review. A gap is a build and task combination with fewer valid rows than the requested target. The second program, 5. Experiment/1. Harness/scripts/paired_analysis.py, is an exploratory comparison reader. It scans scored cells from run sets 011 onward, retains valid records from the two arms named in its source, pairs them by task, seed, model, prompt variant and retry-feedback mode, and prints success and token comparisons. The third program, 5. Experiment/1. Harness/scripts/power_analysis.py, is a deterministic simulation used to inspect the planned power of the paired cost test. Power is the simulated chance of rejecting a null result at a chosen effect and significance level. It writes no campaign record.
The fourth program, 5. Experiment/10. User Interface/publish_snapshot.py, is the publication boundary. It calls the local server and research-state builder, selects named fields, redacts sensitive values, writes static JSON and application files to a temporary directory, checks that the result contains no local controls or write methods, and copies the checked directory to the requested portal output. The portal is the public read-only site. It reads the snapshot; it does not alter the research records.
The analysis plan is the rule book, not an output of these four programs. A Holm correction is a rule that adjusts significance thresholds across a family of tests. An attrition gate is a rule that can demote a conditional cost comparison when the arms lose different proportions of cells. An allowlist is a named set of fields or files permitted through a boundary. 5. Experiment/6. Metrics/statistical_analysis_plan.md version v005 registers the confirmatory family {HN-S, HN-C} for the LAP − NON comparison, with Holm correction, the success-within-budget endpoint, the gated conditional token endpoint, the 15 percentage point attrition gate, and the cache policy. The claim registry is the named inventory of propositions that may be published. In 5. Experiment/0. Plan/claim_registry.json, HN-S is reliability within the token budget and HN-C is conditional token cost. HS-DOC and HS-STR are registered secondary or descriptive comparisons, while H1b is a sensitivity claim and H6 is exploratory.
The source program currently named paired_analysis.py compares H-LAP-L5P with H-STR, not H-LAP-L5P with H-NON, and its module docstring calls its output exploratory. It therefore cannot by itself publish HN-S or HN-C. This distinction is part of the design: a printed exploratory report is not a registry claim. The plan and registry remain the authority for what a confirmatory number would mean. The provenance chain described in operations/results-integrity/B2-count-provenance.md, section 1, begins with a scorer verdict, passes through the aggregator into 5. Experiment/6. Metrics/cell_factor_matrix.csv, and lets the interface count only rows whose valid_for_primary_analysis field is true. A number on the public site is consequently downstream of the scorer and aggregator gates as well as of the publication checks.
The mechanism is invoked in three ordinary ways. An operator or interface asks the coverage script where more valid observations are needed. An operator runs the paired comparison and the power check while assessing the registered analysis. A release action invokes the snapshot builder after the local report and research records are ready. The outputs are deliberately different. Coverage produces planning text or a draft specification. Paired analysis and power analysis produce standard output. Snapshot publishing produces immutable-looking static files whose immutability is operationally supplied by the absence of local controls and HTTP write verbs, not by a claim that a filesystem administrator cannot replace them.
2. The reader’s map of the owning files
The four owning programs are below. The line ranges are taken from the current files and the corresponding code digest sections. The two normative documents are read by people and by adjacent interface code as specified by their own paths; they are not parsed by the four scripts in this chapter.
2.1 analyze_coverage.py
Lines 1 to 33 contain the module description and command examples. Lines 35 to 53 import the standard library, load task_sequences, establish the repository base and define the condition order. analyze_coverage.py:load_rows, lines 56 to 68, is the CSV reader and validity, task and condition filter. analyze_coverage.py:_sorted_conditions, lines 71 to 73, orders conditions. analyze_coverage.py:coverage, lines 76 to 93, counts condition and task, condition and seed, task and seed, model and condition, and hidden-pass combinations. analyze_coverage.py:grid, lines 96 to 111, formats text. analyze_coverage.py:find_gaps, lines 114 to 138, counts valid rows and records integer seeds that have already been used. analyze_coverage.py:runnable_tasks, lines 141 to 160, reads task manifests and omits manifests that cannot be read or that say the task is not valid for measured runs. analyze_coverage.py:suggest_spec, lines 163 to 268, contains the lowest-unused-seed helper, round-robin budget loop, chained-task expansion, budget refusal note and final ordering. analyze_coverage.py:_and_list, lines 271 to 276, joins names. analyze_coverage.py:main, lines 279 to 375, contains the argument parser, mode dispatch, output writing and exit returns. Lines 378 to 379 are the module entry point.
There is no separate finalisation region. suggest_spec writes the draft at lines 369 to 374, while report, gap and JSON modes finish by printing. The file reads but does not own the composite matrix, task manifests or sequence definitions.
2.2 paired_analysis.py
Lines 1 to 26 are the module description and its explicit exploratory selection rules. Lines 27 to 42 import modules, define the current instrument boundary and arm tuple, and begin the estimator region. paired_analysis.py:hodges_lehmann, lines 44 to 49, paired_analysis.py:cliffs_delta, lines 52 to 56, paired_analysis.py:wilcoxon_p, lines 59 to 82, paired_analysis.py:hodges_lehmann_ci, lines 89 to 113, paired_analysis.py:signed_rank_counts, lines 116 to 125, paired_analysis.py:mcnemar_p, lines 128 to 142, and paired_analysis.py:median, lines 145 to 150, are the statistical helpers. paired_analysis.py:collect, lines 153 to 193, scans scoring summaries, filters and extracts cell and metric fields. paired_analysis.py:resolve_pairs, lines 196 to 227, prefers a common run set and otherwise marks a mixed pair. paired_analysis.py:main, lines 230 to 355, contains argument parsing, collection, pairing, per-pair output, the primary success calculation in this exploratory tool, the attrition calculation, the conditional cost calculation, the imputed companion, the task-level clustering calculation and closing labels. Lines 358 to 359 are the module entry point.
There is no file finalisation region because this program writes no file. Its only durable output is whatever an operator captures from standard output. The digest verifies that the input summaries are scoring_summary.json and optional metrics.json files below run-set cell directories.
2.3 power_analysis.py
Lines 1 to 20 describe the simulation and its relation to the statistical plan. Lines 22 to 33 import random and define the fixed seed, trial count and ten (n, effect) rows. power_analysis.py:null_counts, lines 36 to 46, defines the dynamic-programming null distribution. power_analysis.py:critical_deviation, lines 49 to 65, includes the no-rejection result. power_analysis.py:power, lines 68 to 84, simulates signed-rank studies. power_analysis.py:main, lines 87 to 106, prints the table and the false-positive checks. Lines 109 to 110 are the module entry point. There is no file read, file write, network endpoint or separate finalisation region.
2.4 publish_snapshot.py
Lines 1 to 45 contain the publication contract and its two defenses, the explicit allowlists and the redaction backstop. Lines 47 to 71 import dependencies and define paths and path patterns. Lines 73 to 163 define the redaction keys, public data export names, cell and run-set field allowlists, and public application file allowlist. publish_snapshot.py:public_value, lines 166 to 187, recursively redacts values, and publish_snapshot.py:allowlist_rows, lines 190 to 203, filters row fields. publish_snapshot.py:assert_public_safe, lines 206 to 210, publish_snapshot.py:assert_public_bundle, lines 213 to 224, publish_snapshot.py:assert_no_local_scripts, lines 226 to 258, and publish_snapshot.py:assert_no_writes, lines 260 to 277, are the validation helpers. publish_snapshot.py:_write, lines 279 to 283, publish_snapshot.py:_slug, lines 286 to 287, and publish_snapshot.py:_commit, lines 290 to 295, handle JSON writing, filename cleaning and commit lookup. publish_snapshot.py:publishable_drafts, lines 298 to 313, and publish_snapshot.py:withhold_unapproved_drafts, lines 316 to 332, control draft inclusion. publish_snapshot.py:_copy_public_app, lines 335 to 352, copies the named application files. publish_snapshot.py:build, lines 355 to 447, performs extraction, filtering, writing, manifest creation, validation and replacement. publish_snapshot.py:main, lines 450 to 459, parses its two command-line arguments and calls build. Lines 463 to 464 are the module entry point.
The finalisation region is lines 425 to 447. It records whether timing data is nonempty, writes manifest.json, runs the three bundle checks, removes an existing destination if present, and copies the temporary tree to the output. The manifest carries schema literate-programming.public-evidence.v002, generation time, source commit, report map, timing map and research-state path.
2.5 Normative records and plans
5. Experiment/6. Metrics/statistical_analysis_plan.md is the pre-registered plan. Its opening records schema v005 and the requirement that later changes require a version bump and an Experiment Log entry. Its operative regions are unit and pairing at lines 50 to 106, sample size and power at lines 108 to 198, hypotheses and test mapping at lines 202 to 220, multiplicity at lines 245 to 270, exclusions and the attrition gate at lines 272 to 297, stopping and budget at lines 299 to 323, reporting at lines 352 to 358, and cache policy at lines 360 to 370.
5. Experiment/0. Plan/claim_registry.json is a JSON registry rather than a program. Lines 1 to 14 define schema, question and public scope. Lines 16 to 40 define HN-S and HN-C. Lines 42 to 66 define HS-DOC and HS-STR. Lines 68 to 90 define H1b and H6. Lines 93 to 104 define the pending F-PRIM claim. Its finalisation is the registry status and public flag on each claim, not a generated result.
3. Inputs
The tables distinguish an argument, an environment variable, a file and a network endpoint. A blank category means that the source was checked and has no input of that kind.
3.1 Command-line arguments
| Program | Parser line | Argument | Decision made |
|---|---|---|---|
analyze_coverage.py | 280 to 299 | --matrix | Selects the composite CSV instead of the default matrix. |
analyze_coverage.py | 283 to 295 | --report, --gaps, --json, --all | Selects printed form and whether invalid rows are included. |
analyze_coverage.py | 289 to 292 | --target, --tasks, --conditions | Sets the repetition target and optional filters. |
analyze_coverage.py | 293 to 299 | --suggest, --out, --label, --backend, --model, --token-budget | Requests and names a draft batch specification. |
paired_analysis.py | 231 to 234 | positional run_sets | Selects the root containing run-set directories, defaulting to 5. Run Sets. |
power_analysis.py | none | none | The script has no parser and accepts no arguments. |
publish_snapshot.py | 451 to 458 | --portal, --output | Selects the portal base or directly names the bundle destination. |
3.2 Environment variables
The four owning programs read no environment variables. publish_snapshot.py imports server and research_state, and those modules may have their own configuration, but no such environment read is present in this chapter’s owning source. The absence matters because the snapshot builder’s output selection is controlled by explicit arguments and imported server data, not by an unrecorded analysis environment.
3.3 Files read
| Program | File or path | Read at | Decision made |
|---|---|---|---|
analyze_coverage.py | 6. Metrics/cell_factor_matrix.csv or --matrix path | 56 to 60 | Supplies rows for coverage and gaps. |
analyze_coverage.py | 4. Task Library/*/*/task_manifest.json | 151 to 159 | Establishes the live task set and valid_for_measured_runs status. |
analyze_coverage.py | task_sequences.py | 44 to 45 and 185 to 187 | Supplies chain loading and cell ordering. |
paired_analysis.py | 0*/cells/*/scoring_summary.json below the run-set root | 156 to 169 | Supplies run set identity, arm, task, seed and verdict, subject to instrument and validity filters. |
paired_analysis.py | sibling metrics.json | 173 to 185 | Supplies cumulative spend, model identity, prompt variant and retry-feedback mode when present. |
power_analysis.py | none | none | All inputs are constants in the source. |
publish_snapshot.py | server data returned by server.cells, coverage, runsets, reports, timeline, timings and report | 361 to 372 and 406 | Supplies measured rows, summaries, reports and timing detail. |
publish_snapshot.py | research state from research_state.build() | 391 | Supplies the public research-state export. |
publish_snapshot.py | app/ files named by PUBLIC_APP_FILES | 340 to 352 | Supplies the static public application. |
publish_snapshot.py | publishable_drafts.txt beside the script | 307 to 313 | Selects run sets whose drafted analyses may be included. |
publish_snapshot.py | repository path through server.BASE | 292 to 295 | Supplies the short commit for the manifest. |
3.4 Network endpoints
The owning programs define no network client function and call no endpoint. The snapshot builder calls imported Python functions in the local server, not HTTP endpoints. The resulting public application is checked for write verbs at lines 260 to 277, but that check is a scan of JavaScript text, not a network call made during publication.
4. The happy path in order
The complete path has two branches before publication. Coverage is a planning branch and does not need to run before paired analysis. Power analysis is a design validation branch and reads constants rather than the campaign. Publication is the release branch. The functions and ranges below name where each step begins and ends.
4.1 Coverage row loading
load_rows in 5. Experiment/1. Harness/scripts/analyze_coverage.py begins at lines 56 to 58 and ends at lines 59 to 68. It opens the selected CSV with a dictionary reader, keeps rows whose valid_for_primary_analysis value is true unless --all was supplied, then applies task and condition filters. The design reason is in the module description at lines 4 to 7 and the option description at lines 11 to 19: the composite table is not sufficient by inspection to show where valid coverage is thin, so the script makes the population and its filters explicit.
4.2 Coverage aggregation
coverage begins at lines 76 to 83 and ends at lines 84 to 93. It increments dictionaries keyed by condition and task, condition and seed, task and seed, model and condition, and hidden-pass condition and task. The reason is that one grid cannot answer every planning question. The module description identifies the same coverage views at lines 11 to 15, while the loop at lines 84 to 92 keeps the hidden-pass count as a separate descriptive count rather than changing the cell population.
4.3 Gap identification
find_gaps begins at lines 114 to 116 and ends at lines 117 to 138. It counts rows by condition and task, tries to record integer seeds, forms every requested condition and task combination below the target, and sorts the result by sparsity and stable names. The comments and module description say why: a draft should fill the sparsest build and task combinations first, and the lowest unused integer seed gives a deterministic next repetition. A malformed seed is ignored at lines 122 to 125 rather than stopping the entire report.
4.4 Live-task selection and draft construction
runnable_tasks begins at lines 141 to 145 and ends at lines 151 to 160. It walks task manifests, skips an unreadable or invalid manifest, and returns the live task names. The function comment says that historical tasks have data but no live manifest, so a historical gap must not become a new cell. suggest_spec begins at lines 163 to 165 and ends at lines 253 to 268. It loads chain definitions, round-robins through gaps while the cell budget remains, chooses a lowest unused seed, expands a chained task to its required prefix, records a note when the chain cannot fit, orders the final cells, and constructs the draft object. The comments at lines 166 to 183 and 210 to 229 give the design reason: a later chain step cannot be drafted alone because it would not have the predecessor tree it requires, and real predecessor cells count against the budget.
main begins at lines 279 to 300 and ends at lines 302 to 375. It parses the arguments, calls the reader and aggregators, chooses report, gap, JSON or suggestion output, and writes a suggestion at lines 369 to 374. The source comment at lines 350 to 354 says that suggestions use the live roster rather than the historical observed roster. A draft is therefore a proposal for review, not a release or a claim.
4.5 Paired record collection
collect in 5. Experiment/1. Harness/scripts/paired_analysis.py begins at lines 153 to 155 and ends at lines 156 to 193. It scans sorted scoring summaries, parses the run-set number, keeps only run sets 011 and later, splits the cell name into arm, task, phase and seed, keeps only the two ARMS values, reads the scoring summary, rejects a record with invalid_reasons, and supplements it from metrics.json when present. The module comments at lines 7 to 25 give the reason for every major choice. Earlier run sets used a different instrument, invalid cells measured nothing, and a latest record must not be counted twice. The current code groups records by task, seed, exact model identity where present, prompt variant and retry-feedback mode so that unlike instrument choices are not silently pooled.
4.6 Pair resolution
resolve_pairs begins at lines 196 to 199 and ends at lines 211 to 227. It finds run sets common to both arms, chooses the newest common run set when one exists, and otherwise chooses the newest record for each arm while marking the result as mixed. The comment at lines 197 to 209 supplies the reason: an independently newest record can put a change of instrument inside a comparison, while a same-run-set pair keeps the two arms under the same run-set conditions. A mixed pair remains visible in the final count at lines 268 to 269 and in the closing output at lines 352 to 355.
4.7 Exploratory endpoint calculation
main begins at lines 230 to 234 and runs to lines 352 to 355. It retains only keys containing both arms at lines 236 to 240, prints per-stratum and per-pair summaries at lines 241 to 266, and computes the success endpoint at lines 270 to 289. The paired success comparison uses discordant pairs in mcnemar_p, which begins at lines 128 to 136 and ends at lines 137 to 142. The source comment says that equal outcomes carry no information about an arm difference, so only the splits enter the exact test.
The attrition calculation begins at lines 291 to 296. It compares failure counts and labels the cost endpoint as retaining standing or as demoted when the imbalance exceeds 15 percentage points. The conditional cost branch begins at lines 298 to 310. It runs only when both arms have succeeded and token values exist, then calls hodges_lehmann at lines 44 to 49, cliffs_delta at lines 52 to 56, wilcoxon_p at lines 59 to 82 and hodges_lehmann_ci at lines 89 to 113. The code comments explain that a failed cell’s token count is a ceiling rather than the cost of successful work, so the complete-case cost endpoint is conditional and separately gated.
The mandatory imputation branch begins at lines 311 to 332. It assigns TOKEN_BUDGET, 2,500,000 at lines 86 to 87, to a failed or missing-token cell and computes a paired difference for every slot. The comment says why the companion is required: conditioning on two successes removes the reliability information carried by a failed cell. The clustering branch begins at lines 334 to 350 and takes task medians before a signed-rank test. Its comments state that several pairs from one task are not independent observations. The closing lines 352 to 355 label the result exploratory and below the confirmatory minimum.
4.8 Power simulation
main in 5. Experiment/1. Harness/scripts/power_analysis.py begins at lines 87 to 98. It creates a random generator with SEED 20260818, prints the method and both significance columns, iterates through ROWS, and then prints false-positive checks for ten and twenty pairs at effect zero. The comments at lines 6 to 19 explain why simulation is used and why alpha 0.025 matters as the worst case for the second endpoint in a Holm-corrected family.
For each row, power begins at lines 68 to 73 and ends at lines 74 to 84. It calls critical_deviation, which begins at lines 49 to 54 and ends at lines 55 to 65, and that function calls null_counts, which begins at lines 36 to 37 and ends at lines 38 to 46. Dynamic programming computes the signed-rank null distribution once per sample size, then the simulation draws 20,000 normal deltas and counts rejections. When no rejection is possible, critical_deviation returns None and power returns 0.0 at lines 74 to 76. This is an honest small-sample result, not an exception.
4.9 Public application staging
build in 5. Experiment/10. User Interface/publish_snapshot.py begins at lines 355 to 358 and ends at lines 445 to 447. It creates a temporary directory and calls _copy_public_app, which begins at lines 335 to 339 and ends at lines 340 to 352. Only names in PUBLIC_APP_FILES are copied. The module description at lines 16 to 28 gives the reason for default-deny copying: a future local control cannot enter merely because it exists in the local application.
The data calls and public selection begin at lines 360 to 392. The builder obtains cells, coverage, run sets, reports, timeline, timings and research state, keeps only selected coverage keys, applies allowlist_rows to cell and run-set rows, and constructs the export dictionary. The assertion at lines 393 to 398 ensures that every export name belongs to PUBLIC_EXPORTS, then _write writes the JSON. _write begins at lines 279 to 281 and ends at lines 282 to 283. It applies public_value, checks for private absolute paths and writes formatted JSON. public_value begins at lines 166 to 174 and ends at lines 175 to 187, replacing sensitive keys and private paths. The design reason is the two-layer boundary described at lines 30 to 38: explicit field selection is primary and redaction is an independent backstop.
4.10 Report, timing and manifest finalisation
The report loop begins at lines 400 to 409. It reads approved run IDs through publishable_drafts, which begins at lines 298 to 305 and ends at lines 307 to 313, and applies withhold_unapproved_drafts, which begins at lines 316 to 323 and ends at lines 324 to 332. An unapproved drafted analysis is removed from the report and counted as withheld. The timing loop begins at lines 415 to 423 and writes one detail file for each timing cell with both a run ID and cell ID.
Manifest finalisation begins at lines 425 to 441. It records whether the timing summary has cells and writes the schema, timestamp, source commit from _commit at lines 290 to 295, and the report and timing maps. Validation begins at line 442 and continues through lines 443 to 444 with assert_public_bundle, assert_no_local_scripts and assert_no_writes. Only after those checks does the builder remove an existing destination and copy the temporary directory at lines 445 to 447. The reason is transactional staging: an unsafe or incomplete temporary bundle does not replace the previous destination.
5. The state machine
The flow-model file was searched for chapter value 12 and contains no chapter 12 nodes or edges. Its chapter inventory ends with chapter 10, and its stated model version is flow_model.v001. Therefore this state diagram is a hand-authored rendering of the verified source functions, not a rendering of a chapter 12 model subgraph. The model node identifiers for this chapter are none. The record states shown here are the observable states the mechanism can leave in standard output, a draft, a temporary bundle, or the final bundle.
stateDiagram-v2 [*] --> InputsLoaded InputsLoaded --> CoverageReported : analyze_coverage.py 304..341 InputsLoaded --> GapListReported : analyze_coverage.py 343..348 InputsLoaded --> DraftBuilt : analyze_coverage.py 350..374 InputsLoaded --> ExploratoryReport : paired_analysis.py 234..355 InputsLoaded --> PowerReport : power_analysis.py 87..106 InputsLoaded --> TemporaryBundle : publish_snapshot.py 355..398 TemporaryBundle --> DraftsWithheld : publish_snapshot.py 400..413 TemporaryBundle --> ManifestWritten : publish_snapshot.py 415..441 ManifestWritten --> PublicationRefused : publish_snapshot.py 442..444 ManifestWritten --> PublishedBundle : publish_snapshot.py 445..447 CoverageReported --> [*] GapListReported --> [*] DraftBuilt --> [*] ExploratoryReport --> [*] PowerReport --> [*] DraftsWithheld --> ManifestWritten PublicationRefused --> [*] PublishedBundle --> [*]
CoverageReported, GapListReported, ExploratoryReport and PowerReport are terminal output states, not records that the publisher consumes automatically. DraftBuilt is a JSON proposal. TemporaryBundle is the unvalidated staging state. DraftsWithheld is a publication state in which unapproved drafted analysis has been removed from an individual report. ManifestWritten is still temporary until the checks pass. PublicationRefused is an exception state and leaves no new final bundle through this build call. PublishedBundle is the directory copied to the output path.
6. The sequence of one unit of work
A unit of work here is one publication build, because analysis reports are printed rather than handed from one program to the next. The diagram names programs, imported modules, files and the external reader. It is hand-authored because the flow model has no chapter 12 subgraph and therefore supplies no node identifiers for this sequence.
sequenceDiagram participant Operator participant Publisher as publish_snapshot.py participant Server as server module participant State as research_state module participant Source as local records and app files participant Temp as temporary bundle participant Portal as public portal Operator->>Publisher: main(--portal or --output), lines 450..458 Publisher->>Temp: create temporary directory, lines 355..358 Publisher->>Source: copy allowlisted app files, lines 335..352 Publisher->>Server: cells, coverage, runsets, reports, timeline, timings, lines 361..372 Publisher->>State: build research state, line 391 Publisher->>Source: read publishable_drafts.txt, lines 307..313 Publisher->>Temp: write filtered JSON and report files, lines 374..423 Publisher->>Temp: write manifest with commit and maps, lines 433..441 Publisher->>Temp: scan paths, local controls and write verbs, lines 442..444 Publisher->>Portal: replace output with checked tree, lines 445..447 Portal-->>Operator: read static HTML and JSON files
The server and research-state calls are local program calls, not network endpoints. The portal is an external reader only in the sense that it is outside the publishing process and consumes the copied files. It has no participant action that writes back into the experiment.
7. The records
The canonical record rows below belong to this chapter because the analysis and publication programs create or transform them. A reader listed as a program or file means that the source or adjacent interface reads the record there. The paired and power reports are standard output rather than files, so their readers are operators or a separately captured report process.
| Record | Writer and line | Fields or contents | Readers |
|---|---|---|---|
| Coverage report | analyze_coverage.py main, lines 323 to 341 | Row count, validity mode, coverage grids for condition and task, condition and seed, and model and condition, with hidden-pass counts | Operator; results interface imports the same coverage functions as described by B2 section 1, but does not read this stdout. |
| Gap report | analyze_coverage.py main, lines 343 to 348 | Condition, task, current count, needed count and used seeds | Operator; no downstream file reader. |
| Draft batch specification | analyze_coverage.py main, lines 350 to 374, using suggest_spec lines 163 to 268 | label, project, agent_backend, cells, optional model, optional token_budget, and optional sequence_notes | Batch driver when an operator commits the reviewed draft; operator. |
| Exploratory paired report | paired_analysis.py main, lines 230 to 355 | Slot identity, arm verdicts, attempts, tokens, deltas, success counts, McNemar p, attrition label, conditional cost estimates, imputed companion, task medians and mixed-pair count | Operator and any process that captures standard output; no publisher reader is implemented in this source. |
| Power report | power_analysis.py main, lines 87 to 106 | Pair count, simulated effect, power at 0.05 and 0.025, and false-positive checks at zero effect | Operator and the statistical plan, which records the table at lines 140 to 169; no publisher reader. |
| Public cells export | publish_snapshot.py build, lines 374 to 378, then _write, lines 279 to 283 | Generated source, row count and cell rows limited by CELL_FIELD_ALLOWLIST, then redacted | Public client and public application files in the copied bundle. |
| Public coverage export | publish_snapshot.py build, lines 363 to 368 and 384 to 398 | condition_task, sequence_groups, target and valid_only from local coverage | Public client and coverage views. |
| Public run-set export | publish_snapshot.py build, lines 379 to 382 | Run index rows limited by RUNSET_FIELD_ALLOWLIST | Public client and report views. |
| Public reports, timeline, timings and research state | publish_snapshot.py build, lines 384 to 398 | Server reports, timeline, timing data and research_state.build() result after _write redaction | Public client and public application modules. |
| Individual report JSON | publish_snapshot.py build, lines 400 to 409, using withhold_unapproved_drafts lines 316 to 332 | One report keyed by a slug, with drafted analysis removed when its run ID is not approved | Public report view through the report map. |
| Timing-cell JSON | publish_snapshot.py build, lines 415 to 423 | One cell timing result keyed by an index, run ID and cell ID | Public timing view through the timing map. |
| Public application files | publish_snapshot.py _copy_public_app, lines 335 to 352 | The named HTML, JavaScript, CSS and vendor files | Browser loading the published portal. |
| Public manifest | publish_snapshot.py build, lines 425 to 441 | Schema, generation time, source commit, report map, timing map, timing completeness and research-state path | Public client and a reader checking publication provenance. |
The source of the first row is the campaign matrix, not the coverage report. B2 section 1 records the upstream lineage: score_cell.py writes scoring_summary.json, the closeout aggregator writes cell_factor_matrix.csv, and analyze_coverage.load_rows keeps valid rows by the matrix validity field. The publisher does not read the matrix directly in this source. It asks server.cells() for the local server representation and then applies a second public field allowlist.
flowchart LR Cell["cell directory"] --> Score["scoring_summary.json\nscore_cell.py"] Cell --> Metrics["metrics.json\ncell runner"] Score --> Aggregate["aggregate_metrics.py\nvalidity gate"] Metrics --> Aggregate Aggregate --> Matrix["cell_factor_matrix.csv"] Matrix --> Coverage["analyze_coverage.py\nload_rows and coverage"] Matrix --> Server["server.cells and server.coverage"] Server --> Allow["publish_snapshot.py\nallowlists and redaction"] Registry["claim_registry.json"] --> Claim["registered claim name"] Plan["statistical_analysis_plan.md"] --> Decision["endpoint, gate and cache policy"] Decision --> Claim Allow --> Snapshot["snapshot JSON and manifest"] Claim --> Snapshot Snapshot --> Portal["read-only public portal"]
This lineage diagram is hand-authored from B2 section 1, the plan, the registry, and publish_snapshot.py lines 361 to 447. It is not generated from the flow model. The registry names the claim and the plan defines the decision rule, but the current publisher source does not show a direct call that reads claim_registry.json; the registry relationship is therefore a publication requirement and provenance link, not an asserted runtime read by publish_snapshot.py.
8. The loops and the waits
analyze_coverage.py has a for loop over rows in coverage at lines 84 to 92. Its bound is the number of rows returned by load_rows; it stops after every selected row has contributed to each applicable count. find_gaps has a loop over selected rows at lines 119 to 125 and nested loops over all conditions and tasks at lines 130 to 138. Their bounds are the selected row count and the Cartesian set of requested conditions and tasks. They stop after all counts and all combinations are considered.
suggest_spec has a while loop at lines 202 to 231. Its bound is not a fixed integer because n_cells is an argument, but the number of appended cells cannot exceed n_cells; it stops when the cell list reaches the requested budget or every remaining gap has no need. Its nested for loop at lines 203 to 208 is bounded by the number of gaps on each round and stops early when the budget is full. The _next_seed loop at lines 192 to 197 increments from one until it finds an integer not in that gap’s used-seed set. There is no sleep and no retry. A chain that cannot fit is marked complete for drafting at lines 221 to 231 and is not half added.
paired_analysis.collect loops over the sorted path result at lines 156 to 169. Its bound is the number of matching scoring_summary.json paths under the root, and its stop condition is exhaustion of that iterator. resolve_pairs loops over grouped keys at lines 212 to 226 and over each arm’s records when choosing a common or mixed record at lines 217 to 224. The main report loops over strata at lines 240 to 245, pairs at lines 252 to 266, all pairs for the imputed companion at lines 318 to 323, and task groups at lines 339 to 344. Every loop stops after its finite collection is exhausted. There is no wait, backoff or retry policy. A missing pair is omitted rather than retried.
power_analysis.main loops over the ten source rows at lines 95 to 98. power loops over TRIALS, exactly 20,000, at lines 77 to 84 for each row and significance level. null_counts and signed_rank_counts loop over ranks from one through n at lines 39 to 45 and 118 to 124 in the paired program. These loops are finite and have no wait or retry. The fixed random seed makes the simulation repeatable within the source’s random implementation.
publish_snapshot.build loops over reports at lines 403 to 409 and timing cells at lines 416 to 423. Their bounds are the lengths of reports.get('run_sets', []) and timings.get('cells', []), with missing run or cell identifiers skipped at lines 418 to 420. The validation loops scan every file at lines 247 to 258, every JSON file at lines 222 to 224, and every JavaScript file at lines 267 to 276. There is no wait or retry. A check raises rather than asking for another source or attempting a repair. The only replacement operation is the final removal and copy of the destination after validation at lines 445 to 447; it is not a retry.
9. Guards and refusals
| Guard location | Check | Refusal or alteration |
|---|---|---|
analyze_coverage.py lines 61 to 63 | Validity filter | Excludes rows whose validity field is not true unless --all is requested. |
analyze_coverage.py lines 122 to 125 | Integer seed conversion | Ignores an invalid seed for seed uniqueness while continuing the report. |
analyze_coverage.py lines 154 to 158 | Manifest parsing and measured-run flag | Omits an unreadable, invalid JSON or explicitly unfit task from drafting. |
analyze_coverage.py lines 202 to 231 | Remaining cell budget for a chain | Skips a later chain gap and adds a sequence note instead of drafting a partial chain. |
analyze_coverage.py lines 304 to 308 | Empty selected row set | Prints an error and returns one. |
analyze_coverage.py lines 355 to 361 | Empty live task set in suggest mode | Prints an error and returns one. |
paired_analysis.py lines 159 to 160 | Run set before 011 | Omits the cell because the instrument differs. |
paired_analysis.py lines 163 to 170 | Cell name and validity | Omits malformed names and records with invalid reasons. |
paired_analysis.py lines 166 to 167 | Arm membership | Omits arms outside the two configured values. |
paired_analysis.py lines 213 to 225 | Pair completeness and same-run preference | Reports only keys with both arms; uses a mixed marker when no common run set exists. |
paired_analysis.py lines 98 to 100 | At least six deltas for a confidence interval | Returns (None, None) rather than inventing an interval. |
paired_analysis.py lines 291 to 296 | Failure imbalance over 15 percentage points | Demotes the conditional cost output to description. |
paired_analysis.py lines 258 to 262 | Both arms passed and both token values exist | Refuses to add a conditional cost delta for a failed or missing-token pair. |
power_analysis.py lines 59 to 65 and 74 to 76 | No critical rejection deviation | Returns power 0.0 rather than raising. |
publish_snapshot.py lines 200 to 203 | Data field allowlists | Drops unlisted cell and run-set fields. |
publish_snapshot.py lines 175 to 187 | Redacted key and path handling | Replaces sensitive values and removes or marks private paths. |
publish_snapshot.py lines 206 to 210 | Private absolute path | Raises ValueError. |
publish_snapshot.py lines 213 to 224 | Local config or unsafe JSON | Raises ValueError. |
publish_snapshot.py lines 226 to 258 | Local scripts and controls | Raises ValueError on a forbidden reference. |
publish_snapshot.py lines 260 to 277 | HTTP write verb | Raises ValueError on POST, PUT, DELETE or PATCH text in project JavaScript. |
publish_snapshot.py lines 393 to 395 | Export name outside PUBLIC_EXPORTS | Raises AssertionError. |
publish_snapshot.py lines 307 to 313 | Missing approval list | Treats the approved set as empty, so drafted analyses are withheld. |
publish_snapshot.py lines 445 to 447 | Final destination replacement | Occurs only after all checks pass. |
The decision tree below makes the two analysis gates visible. It is hand-authored from the statistical plan lines 247 to 270 and 272 to 297, the paired program lines 291 to 332, and the publisher checks at lines 442 to 444. The registry identifiers are the claim nodes that a reviewed result would have to name. The tree is not a claim that this current exploratory paired script computes HN-S or HN-C, because its configured arms are different.
flowchart TD Start["eligible records"] --> Valid{"valid_for_primary_analysis?"} Valid -- no --> Attrition["retain in attrition report\nexclude from statistics"] Valid -- yes --> Pair{"both arms and matched seed?"} Pair -- no --> NoPair["no paired claim"] Pair -- yes --> Success["HN-S success within budget\nall eligible cells"] Success --> Gate{"failure imbalance <= 15 percentage points?"} Gate -- no --> Demote["HN-C descriptive only"] Gate -- yes --> Cost["HN-C conditional token cost\nboth-success pairs"] Cost --> Cache{"cache policies agree in sign?"} Cache -- no --> Dependent["cache-policy-dependent claim"] Cache -- yes --> Registry["registry claim may be reported"] Registry --> Publish{"public bundle checks pass?"} Publish -- no --> Refuse["publication refused"] Publish -- yes --> Portal["number on read-only portal"]
10. Every unhappy path
The table uses four parts for every path. Each entry states what the step is supposed to do, why it works that way, what triggers the failure and what record it leaves, and what it costs. Where the source does not set a process exit code, this chapter says so rather than inferring one. The downstream columns distinguish the batch driver, scorer, aggregator and interface. None of these analysis failures runs the scorer or changes an already written cell verdict.
| Trigger and path | What the step is supposed to do and why | Trigger, record and status | Cost and downstream effect |
|---|---|---|---|
No rows match filters, analyze_coverage.py lines 304 to 308 | load_rows should provide the selected population. The filter is applied first so a report cannot silently count out-of-scope rows. | An empty result after validity, task or condition filters triggers the branch. It writes the stderr text no rows matched, leaves no data record, returns exit code 1 and does not produce a report. | No driver action, no scorer action and no aggregator action. The interface receives no report from this invocation. |
| No live registered tasks in suggest mode, lines 350 to 361 | runnable_tasks should provide tasks that can actually be staged. A live manifest prevents historical gaps becoming new work. | No live task survives the filters. The script writes an stderr refusal, leaves no draft, and returns exit code 1. | The driver has no draft to launch, the scorer and aggregator are untouched, and the interface can show only the refusal if it captures stderr. |
| Invalid task manifest, lines 151 to 159 | runnable_tasks should inspect each manifest. Skipping a bad manifest keeps one unreadable task from hiding other usable tasks. | OSError or JSONDecodeError causes a silent skip. No per-task error record is written and the overall status remains whatever the remaining tasks produce. | Coverage may be incomplete. No driver, scorer or aggregator change occurs. The interface can show a smaller draft or report without knowing which skipped manifest caused it. |
| Invalid seed, lines 122 to 125 | find_gaps should preserve the set of used integer seeds. Ignoring a noninteger avoids aborting a whole coverage report over one malformed row. | ValueError or TypeError causes that seed to be omitted from the used set. No error record is written and the process continues with a possibly altered gap. | A later suggestion may select a seed already represented by the malformed row. The driver, scorer and aggregator are unchanged, and the interface sees only the resulting gap output. |
Chain budget overflow, suggest_spec lines 210 to 231 | The suggestion should fill a gap without exceeding its requested cell budget. A chain is expanded from its first step so a partial dependency is not drafted. | The required prefix would not fit. The gap is marked no longer needed for this draft and sequence_notes records the omitted task and required prefix. The process remains successful with the incomplete draft. | The driver must review or revise the draft. No cell is scored or aggregated, and the interface can display the explanatory note. |
Run set below 011, paired_analysis.py lines 156 to 160 | collect should compare a common instrument. The current-instrument boundary prevents pilot and repaired instrumentation from being pooled. | The run-set number is below 011. The summary is skipped with no new record and no failure status. | The exploratory sample shrinks. The driver, scorer and aggregator are untouched, and the interface cannot distinguish this omission unless it reads the printed counts and source records. |
| Cell name format mismatch, lines 161 to 165 | collect should recover arm, task, phase and seed for pairing. A strict split prevents a guessed identity. | cell.split('__') does not yield four values. The cell is skipped with no output record and no source-defined exit code. | The pair may disappear. No driver, scorer or aggregator change occurs, and the interface sees no cell-level reason in this report. |
Arm outside ARMS, lines 163 to 167 | collect should include only the configured comparison. The arm guard prevents an unrelated condition entering the paired table. | The parsed arm is not H-LAP-L5P or H-STR. The cell is skipped with no new record and no source-defined exit code. | The comparison remains restricted but may have fewer slots. Other pipeline programs are unaffected. |
| Invalid reasons in scoring summary, lines 168 to 170 | collect should analyze valid cells only. A nonempty invalid-reason list means that the cell did not produce an admissible measurement. | The summary has invalid_reasons. The cell is dropped and no paired record is emitted. The process continues with no source-defined failure exit code. | The paired sample and power interpretation are reduced. The driver, scorer and aggregator are not changed, while the interface retains the invalid source record through its own data paths. |
No matched pair, main lines 236 to 240 | The report should use paired observations so task difficulty is controlled by matching. Requiring both arms avoids an unpaired arm total being read as a paired contrast. | A grouped key lacks one configured arm. The key is excluded from pairs; no error record or nonzero status is produced. | No arm comparison is printed for that key. The driver, scorer and aggregator are unaffected, and the interface receives no paired result from this stdout-only tool. |
Fewer than six deltas, hodges_lehmann_ci lines 98 to 113 | The function should return a confidence interval only where its exact inversion has enough observations. Returning null bounds exposes insufficiency. | len(deltas) < 6 returns (None, None). The paired report prints the text indicating too few pairs and leaves no file record. | The point estimate and other statistics may remain, but the interval is absent. No driver, scorer or aggregator effect exists, and the interface cannot publish the absent interval through this program. |
No both-success deltas, main lines 298 to 310 | The conditional cost endpoint should be computed only where both arms succeeded and token values exist. This avoids treating a failure ceiling as task cost. | deltas is empty, so the secondary block is not printed. No cost result record is written and no exit code is set by the source. | The cost comparison is unavailable while success counts may remain. The driver, scorer and aggregator are unaffected, and the interface cannot infer a cost result. |
Attrition imbalance above 15 percentage points, main lines 291 to 296 | The gate should protect the conditional endpoint from treatment-correlated selection. The source labels the endpoint rather than hiding the surviving pairs. | The absolute failure-rate difference exceeds 15 percentage points. Standard output records OPEN, cost endpoint is demoted to description; no process failure is set. | HN-C style interpretation loses confirmatory standing for this output. The driver, scorer and aggregator do not change, and the interface must show the demotion rather than a confirmatory number. |
No critical deviation, power_analysis.py lines 49 to 65 and 74 to 76 | critical_deviation should identify a rejection boundary. Returning no boundary is the honest result for a small sample or strict alpha. | No rejection is possible, so crit is None and power returns 0.0. Standard output carries the zero in the table and no error status is set. | The design appears unpowered at that row. No driver, scorer or aggregator action follows, and the interface has no automatic consumer. |
Missing publishable_drafts.txt, publish_snapshot.py lines 307 to 313 | publishable_drafts should provide the approved run-set set. Empty approval is safer than publishing unread analysis. | An OSError returns an empty set. No error record is written; reports are processed with all drafted places unapproved. | Draft analysis is withheld, but measured report data can still publish. The driver, scorer and aggregator are unaffected; the interface sees the withheld-draft count where one was removed. |
Private absolute path, assert_public_safe lines 206 to 210 | The bundle should contain repository-relative or public values only. The check is a second defense after redaction. | A matching private path raises ValueError. The temporary bundle is not copied to the final output, and the source does not catch the exception, so the command exits with the interpreter’s failure status rather than a declared numeric code. | Publication stops. No driver, scorer or aggregator action occurs, and the interface keeps the prior destination if one existed. |
Local configuration in snapshot, assert_public_bundle lines 213 to 224 | The bundle should contain evidence, not local experiment controls. The scan checks both forbidden config presence and JSON safety. | A snapshot/config path or unsafe JSON raises ValueError. No final replacement occurs and no explicit exit code is set by the source. | Publication fails before replacement. The driver, scorer and aggregator are unaffected; the interface is not refreshed. |
Local control reference, assert_no_local_scripts lines 226 to 258 | Public HTML and JavaScript should not carry local controls or endpoints. A text scan catches accidental copies. | A forbidden script, control ID or local endpoint raises ValueError. The temporary bundle remains unpromoted and no explicit exit code is set. | The public copy is not produced. The interface remains on the previous checked bundle, with no effect on execution or scoring. |
HTTP write verb, assert_no_writes lines 260 to 277 | The public application should be read-only. Scanning project JavaScript catches a write capability even if a visual control is hidden. | A project script contains a quoted POST, PUT, DELETE or PATCH verb. ValueError stops the build, without a final bundle or source-defined exit code. | Publication is refused. The driver, scorer and aggregator are unchanged and the interface is not refreshed. |
Export outside allowlist, build lines 384 to 398 | build should write only named public exports. The assertion catches future code that adds an unapproved data product. | set(exports) exceeds PUBLIC_EXPORTS. AssertionError stops the build before final copy, without a source-defined exit code. | No new public bundle is installed. The other pipeline programs are unaffected and the interface keeps its previous state. |
Git commit lookup failure, _commit lines 290 to 295 | _commit should tie the manifest to a short source revision. Returning a sentinel preserves the rest of the evidence while disclosing the missing revision. | subprocess.run returns a nonzero status or cannot provide stdout under the function’s nonraising check=False call. The manifest records source_commit: unknown; the build otherwise continues. | Provenance is weaker but publication can proceed. No driver, scorer or aggregator effect exists; the interface shows the sentinel in the manifest. |
Missing report identifiers or cell identifiers, build lines 403 to 423 | The builder should create addressable individual report and timing files. Skipping incomplete entries avoids inventing filenames. | A report entry without a run ID is not expected by the direct indexing at line 404 and may raise a normal exception; a timing entry without run ID or cell ID is skipped at lines 418 to 420. The source does not define a status or exit code for either case. | A malformed report can abort publication, while a malformed timing row is absent from the detail map. Other pipeline programs do not change and the interface either remains old or sees an incomplete timing map after a successful build. |
11. The metrics produced or fed
The primary population measure is valid_for_primary_analysis. B2 section 1 records that the scorer’s valid_for_primary_analysis decision is copied unchanged by the aggregator into 5. Experiment/6. Metrics/cell_factor_matrix.csv, and analyze_coverage.load_rows reads it at lines 61 to 63. Coverage counts therefore become condition and task counts, condition and seed counts, task and seed counts, model and condition counts, and hidden-pass counts through coverage lines 76 to 93. These are planning and descriptive measures. They do not become HN-S or HN-C merely because they appear in a grid.
The registered success endpoint is defined by the statistical plan at lines 207 to 210 as success within budget, with the recorded-field formula including visible pass, hidden pass, no contamination, no more than five passes and spend no greater than 2,500,000. The current paired_analysis.py success output at lines 275 to 289 uses the score_states.overall verdict from each scoring summary and counts passes per arm. That output is exploratory and is configured for H-LAP-L5P versus H-STR; it is not a verified HN-S result for the registry’s H-NON comparator.
The conditional cost measure is tokens_to_success, represented in the paired tool by metrics.json field cumulative_spend at lines 173 to 187 and reduced to paired differences at lines 258 to 263. hodges_lehmann, wilcoxon_p, cliffs_delta and hodges_lehmann_ci turn those differences into a shift estimate, exact p-value, effect size and interval. The plan requires the complete-case cost result only while the 15 percentage point gate is closed and requires the imputed companion. The tool’s companion at lines 317 to 332 charges failed cells the TOKEN_BUDGET ceiling. Both are advisory in this implementation because the tool labels its report exploratory and does not write the registered paired_response_matrix.csv.
The power table is advisory design evidence. power_analysis.py prints simulated power at two alpha levels and false-positive checks. The statistical plan copies the resulting table at lines 152 to 169 and states that the 0.025 column is the relevant Holm worst case. It also records the normal-delta and heavy-tail limitation at lines 191 to 198. The table informs sample-size interpretation; it is not an observed campaign metric.
The public cell columns are the fields in CELL_FIELD_ALLOWLIST, defined at lines 107 to 135 of publish_snapshot.py. They include outcome fields such as visible_pass, hidden_pass, tokens_to_success, valid_for_primary_analysis, exact_model_id, token components, timing components and analysis factors. The publisher does not calculate these fields. It obtains server data at lines 361 to 372, keeps the named columns at lines 374 to 382, and writes them at lines 397 to 398. RUNSET_FIELD_ALLOWLIST similarly supplies the public run-set columns. Reports, timelines, timings and research state are passed through the same _write redaction path. The manifest’s source_commit, schema and maps are publication metadata, not campaign response measures.
The cache policy is a claim guard, not a hidden transformation in this publisher. The plan says at lines 362 to 370 that the dissertation headline excludes cache_read, that the included variant is mandatory, and that a sign disagreement makes the claim cache-policy-dependent. The public allowlist carries both cache token fields and cache_hit_ratio, but this chapter cannot claim that publish_snapshot.py computes or adjudicates the two policies. That decision remains in the analysis and report layer described by the plan.
12. Tests
The test inventory exercises important functions, but the mapping from every unhappy path in section 10 to a test is not made in the repository. The verified tests are as follows. 5. Experiment/1. Harness/scripts/tests/test_analyze_coverage.py imports analyze_coverage and covers validity filtering, gap finding, suggestion construction, command-line modes and runnable task selection at its lines 37 to 105. 5. Experiment/1. Harness/scripts/tests/test_prompt_variant_and_retry_feedback.py imports paired_analysis and covers separation of retry-feedback modes at lines 355 to 358 and the surrounding fixture tests. The source search found no dedicated test for every paired-analysis skip, the six-pair confidence interval refusal, the no-delta branch, the 15 percentage point demotion, or the power function’s no-critical-deviation branch.
The publication tests are more direct. 5. Experiment/10. User Interface/test_phd_publish.py covers public-value path safety at lines 84 to 98, local configuration refusal at lines 107 to 111, a real temporary build from lines 122 to 138, allowlisted export names and absent local exports at lines 211 to 227, manifest maps at lines 246 to 260, and the bundle scans at lines 180 to 194. 5. Experiment/10. User Interface/test_evidence_public.py covers the public allowlist at lines 142 to 153 and the no-local-script and no-write checks at lines 158 to 173. 5. Experiment/10. User Interface/test_timeline_phases.py covers reuse of the same timeline function at lines 181 to 195. 5. Experiment/10. User Interface/test_research_state.py covers current state reconciliation and claim-registry path validation, including validate_registry at lines 57 to 62.
The tests that cover the unhappy paths can therefore be stated precisely. test_analyze_coverage.py covers the empty-row and no-live-task command paths through its main tests, but the source search did not identify a separate assertion for invalid seed handling, invalid manifest JSON, or chain budget overflow. test_phd_publish.py covers local configuration refusal and private path refusal, and its real-build tests cover allowlist and final bundle shape. test_evidence_public.py covers local script and write-verb refusal. The missing approval-list branch, export assertion, git sentinel, malformed report entry, paired skip branches, paired demotion branch, and power no-rejection branch have no named test found by this review. The mapping of tests to exit paths was not made.
13. The dated incidents that shaped the code
The comments in the owning files carry design history that explains why guards exist. They are not errors to remove. paired_analysis.py lines 9 to 12 record the 2026-08-15 instrument ruling that run sets 002 to 009 use a different prompt and run set 010 has no measured work. Lines 19 to 25 explain why pairing is required and why conditional cost excludes failed cells. Lines 38 to 42 state that pre-registered estimators are computed even though the current sample is exploratory. Lines 90 to 97 record the 2026-08-20 confidence-interval provenance problem and the reason the exact inversion is kept in source. Lines 197 to 209 record the mixed-run-set error in which a 20,291-token swing could decide a tie-break, motivating common-run preference. Lines 270 to 275 record that the primary completion endpoint had previously been omitted from the tool. Lines 311 to 316 explain the mandatory imputation companion and its protection against making a frequently failing arm look cheaper. Lines 334 to 350 record the non-independence of repeated pairs from one task. Lines 352 to 355 keep the exploratory and instrument-repair status visible.
analyze_coverage.py lines 50 to 53 record the registered build order. Lines 142 to 149 explain why retired tasks with historical data must not become new drafted work. Lines 166 to 183 record the chained-task incident in which drafts carried a later step without its predecessor, including drafts that were refused at preflight and one that was launched and died. Lines 210 to 229 preserve the budget rule that refuses a partial chain. Lines 350 to 354 explain the live roster rule, which prevents a historical observed roster from determining future work.
power_analysis.py lines 6 to 19 record the correction from a closed-form assumption to simulation and the importance of alpha 0.025. Lines 10 to 14 record the performance reason for dynamic programming. Lines 49 to 54 state that no rejection boundary is an honest result for very small samples. Lines 100 to 106 identify the false-positive check as the validation of the implementation.
publish_snapshot.py lines 4 to 14 record the boundary between the local research console and the public evidence application. Lines 16 to 44 record the allowlist decision and the list of local material that must never enter the bundle. Lines 73 to 83 retain the independent redaction backstop. Lines 85 to 105 explain default-deny field selection. Lines 226 to 231 explain that the local-script scan is a defense against controls crossing the boundary. Lines 260 to 265 explain why the no-write scan is mechanical rather than inferred from visible buttons. Lines 298 to 305 preserve the rule that an unreviewed machine-written draft is not a judgment the campaign stands behind. Lines 316 to 322 preserve the distinction between a missing analysis and a withheld draft. Lines 425 to 431 record the timing completeness marker so the public chart cannot silently present an empty export as a complete result.
The dated plan decisions that shape these guards are in statistical_analysis_plan.md. The v005 re-registration at lines 26 to 48 is dated 2026-08-11 and establishes HN-S, HN-C, the LAP and NON comparison, Holm correction, the attrition gate and the cache-policy decision. The numerical power computation is dated 2026-08-18 at lines 123 to 150. The plan also records the 2.5 million token ceiling and the cache-read exclusion at lines 299 to 370. The chronology corroborates these dates in Appendix A, including the 2026-08-11 v005 row at line 75 and the 2026-08-18 power work in the plan itself. The later paired-tool changes that distinguish model, prompt variant and retry feedback are represented by its source grouping at lines 171 to 192 and by the chronology’s 2026-09-01 entries, but they do not turn its current H-LAP-L5P versus H-STR output into the registry’s confirmatory HN comparison.
14. The weakest claim, what was not checked, and the token line
The weakest claim is that a number printed by the exploratory paired tool can be carried into a named public claim. The source verifies the exploratory calculation and the registry verifies the claim definition, but no direct source in this chapter connects paired_analysis.py output to a registered HN-S or HN-C result record. Not checked: a complete confirmatory implementation for H-LAP-L5P versus H-NON; a generated paired_response_matrix.csv writer in this four-file path; a test mapping for every unhappy path; the exact server data structures returned to publish_snapshot.py; the final portal behavior after copying the bundle; and a chapter 12 subgraph in flow_model.v001.json, which the file does not contain.
Model: openai/gpt-5.6-luna via OpenRouter; tokens: see the job ledger