Written 2026-09-12 by Claude Fable 5.1 to the standard set in 5. Experiment/11. Detailed Design/specifications/C0-detailed-design-specification.md and shown by the reviewed sample chapter 06-execution-of-a-cell.md. Every line number below was read from the source at commit 6e2ee2c2b; the scripts this chapter owns have the same Git blob hashes at that commit as at commit b69ae5977. The local-model draft operations/detailed-design/drafts/chapter_01.md was the lead for the file inventory and nothing else; section 14.1 lists what it got wrong. The flow model 5. Experiment/11. Detailed Design/flow-model/flow_model.v001.json has no nodes for this chapter (its open item 24 says chapters 01, 07, 11, 12 and 13 are outside stage 1), so the diagrams of sections 5 and 6 are drawn from the diagram plan’s sources in the specification’s section 5 and not from a model subgraph; each diagram says so and proposes the node identifiers a later stage of the model could adopt.
Owning files. Six configuration files under 5. Experiment/1. Harness/config/: run_defaults.json (2,275 bytes), model_matrix.json (2,710), scoring_contract.json (5,791), reproducibility_schema.json (2,094), agent_handoff.schema.json (6,308) and hidden_judge.json (7,036). Six prompt templates under 5. Experiment/1. Harness/prompts/: maintenance.single_agent.v006.md (3,045 bytes, 45 lines, the prompt of record), maintenance.single_agent.v005.md (2,588), maintenance.single_agent.v004.md (923), maintenance.single_agent_control.v003.md (750), maintenance.single_agent_lap_selective.v003.md (1,075) and artifact_hint.v001.md (263). Thirteen scripts under 5. Experiment/1. Harness/scripts/: strip_lap.py (270 lines, 9,606 bytes), materialize_feature_bases.py (276 lines, 12,114), instrument_profile.py (397 lines, 17,098), freeze_non.py (149 lines, 6,795), freeze_non_parity_repair.py (146 lines, 6,519), build_phases.py (84 lines, 4,209), lint_hidden_scorers.py (106 lines, 4,338), lint_task_briefs.py (210 lines, 9,342), lint_task_literals.py (218 lines, 9,561), lint_retry_feedback.py (142 lines, 5,791), gen_retry_feedback.py (156 lines, 5,980), decontaminate_briefs.py (150 lines, 7,062) and verify_task_absence.py (179 lines, 8,158). And the schema and guide documents of the three libraries: 2. Project Library/project_manifest_schema.md (3,020 bytes) and README.md (2,666); 3. LAP Profile Library/README.md (6,037), LAP_LEVELS_GUIDE.md (9,315), measurement_policy.md (6,831), traceability_schema.md (7,581) and lap_profile_index.csv (1,444); 4. Task Library/README.md (15,636), task_manifest_schema.md (17,523), seeded_defect_guide.md (6,973), task_scoring_rubric.md (968) and Project H/README.md (17,498). PROBE_AUTHORING.md and the hidden-check convention are chapter 08’s. Files this chapter reads into but does not own: run_cell.py (the cell runner, which reads the configuration and the prompt and calls the brief lint, chapter 06), score_cell.py (the scorer, whose scorer classifier the hidden-scorer lint borrows, chapter 08), preflight_batch.py and analyze_coverage.py (the readers of the task flag, chapters 02 and 12), aggregate_metrics.py (chapter 10), handoff_lib.py (chapter 05), ledger_lib.py (chapter 09, which borrows the stripper) and 4. Task Library/_hidden_lib/fiveb_judge.py (chapter 08).
1. What this chapter is for and where it sits in the life of the experiment
A cell is one measured attempt: one coding agent, one maintenance task, one product build, one repetition seed, run at most five times within one cumulative token budget. Before any cell can run, three things must exist and be trustworthy. The product builds must exist as frozen trees that differ only in their documentation: the documented reference build, its mechanically stripped twin, an independently written control, and the seven intermediate documentation levels between stripped and documented. The tasks must exist as folders whose agent-facing text says nothing about the study, whose hidden checks are executable and honest, and whose registration record says whether a batch may contain them. And the constants every later program reads, the pass ceiling, the token budgets, the model classes, the validity rules, the handoff shape and the judge’s settings, must live in one place each. This chapter owns those three things: the configuration files, the prompts, and the scripts that build and check the libraries.
Nothing in this chapter spends a measured token or writes into a run set. The scripts here run before a batch, by the operator’s hand, and two of them are also called at run time: the runner calls the brief lint before every measured invocation and refuses the cell on a hit (run_cell.py line 2089, chapter 06), and the ledger builder borrows the stripper to decide whether an edit changed only comments (ledger_lib.py lines 47 to 52, chapter 09). The configuration files are read at run time by the runner, the scorer, the aggregator and the preflight; the prompt of record is read by the runner and copied into every run set (run_cell.py 1786 to 1788 and 2222). The reason the constants are files and not code is stated in hidden_judge.json for its own case and holds for all six: a choice made once, in one named file outside any task folder, is visible and can be changed in one place, where a choice buried in whichever script needs it is neither (hidden_judge.json line 6).
The reason the builds are derived rather than built separately is the design lock the instrumenter’s docstring quotes: across the documentation levels of one project the executable code is identical, because each level is a fresh instrumentation of the same clean code, never a separate build, since a language model building each level separately would give each level different code and confound the comparison (instrument_profile.py lines 4 to 13). The pipeline is therefore: build and freeze one reference variant with the full documentation (build_phases.py), derive the clean code by a deterministic strip (strip_lap.py), instrument that clean code to each level with a conformance gate that proves the code stayed byte-identical (instrument_profile.py), and freeze the independently written control beside them (freeze_non.py, freeze_non_parity_repair.py). A task’s feature-removal realisation for the stripped arm is then derived by the same strip, subtract-then-strip, so that it anchors on the stripped tree by construction (materialize_feature_bases.py).
2. The reader’s map of the owning files
2.1 The six configuration files
run_defaults.json (35 lines) carries the pass ceiling max_passes of 5 (line 3), the cumulative token budget of 2,500,000 with its basis (4 to 5), the budget of 30,000,000 for the dgx_claude back end with the measured basis of run set 035 seed 2 (6 to 7), the legacy checkpoint (8), the flakiness rerun count of 3 (9), the cost model the preflight prints (10 to 15), the pass-at-budget grid (16), the correctness gate (17 to 20), the primary metric tokens_to_success and the four secondary metrics (21 to 27), and the conditions per project with the note that Project H’s treatment arm is the eight-rung dose ladder (28 to 33). model_matrix.json (66 lines) pins the measured model class haiku to claude-haiku-4-5-20251001 from the containerised run sets’ transcripts (5 to 12), lists sonnet and opus as future factors (13 to 26), names the fields a run manifest must carry (28 to 36), and records the corpus-build roles and their token treatment (37 to 65). scoring_contract.json (74 lines) names the four score states and the correctness gate (3 to 9), the primary delta per project (10 to 14), the seven validity reasons (15 to 23), and four dated blocks that explain a reason or a step: hidden_scorer_missing (24 to 29), shell_unavailable with the 134 affected cells (30 to 43), the full-suite regression gate (44 to 54), the flakiness rerun (55 to 64) and the parser’s retention rule (65 to 73). reproducibility_schema.json (60 lines) names the fields a task manifest, a run manifest, a reproducibility manifest and a profile manifest must carry (3 to 59). agent_handoff.schema.json (195 lines) is a JSON Schema for the handoff note an agent leaves, HANDOFF.json: ten required members (8 to 19), each interface with its method, path, parameters, examples and the numbered requirements it satisfies (43 to 120), and the rule that it must never name internal functions, classes or modules (5, 45, 68). hidden_judge.json (58 lines) decides whether the fourth level of a hidden check, a language model shown the browser’s evidence, runs at all (8 to 9), which judge does it (11 to 37), what evidence is stored (39 to 43), and four rules: the judge sees behaviour only, runs only when the deterministic levels failed, cannot overturn a deterministic pass, and a judge-only pass is flagged separately (45 to 57).
2.2 The six prompt templates
maintenance.single_agent.v006.md is the one prompt every build has run under since 2026-09-03: an order of nine operations (lines 7 to 25), of which the second reads a predecessor’s HANDOFF.json first and the ninth requires the agent to update any documentation the repository keeps for the parts it changed, in words that act only where such documentation exists; the sentence that the task description is the definition of done (27 to 28); and the handoff instruction with the JSON shape (30 to 45). v005 (2026-09-02) is the same without the second and ninth steps; v004 is the same without the handoff; v003 was two prompts, one for the control that ended by naming the treatment (_control.v003.md line 18) and one for the treatment that opened by listing the artefacts (_lap_selective.v003.md lines 7 to 8), and the runner’s comment at run_cell.py 81 to 103 records why that pair could not attribute a win to documentation rather than to being told to use documentation. artifact_hint.v001.md is the one paragraph the artifact_hint variant appends, with two placeholders the runner fills from the staged copy’s own inventory (lines 7 to 8; run_cell.py 110 to 131 and compose_artifact_hint 869 to 922, chapter 05). The two retired v003 files are kept so that old run records can be read and are excluded from the brief lint by name (lint_task_briefs.py 52 to 60).
2.3 The thirteen scripts
strip_lap.py: sha256_file (36), the docstring and comment strippers _docstring_spans (42 to 71), strip_docstrings (74 to 88), strip_comments (91 to 117) and their composition strip_python_source (120 to 123), the manifest readers load_content_stub_files (126 to 143) and load_lap_files (146 to 156), and main (159 to 266). materialize_feature_bases.py: the exclusion sets (49 to 65), _scratch_ignore (68), _copy_source_tree (76), _strip_tree (84), _tree_hashes (94), verify_strip_matches_frozen (112), _diff_trees (127), _dry_run_clean (141), materialize_task (149 to 205) and main (208 to 272). instrument_profile.py: the profile registry (43 to 55), profile_root (58), compose_prompt (72 to 132), the command-line seam invoke_claude_cli (138 to 151), _next_transcript_path (154), copy_clean_tree (170), the five gate checks check_code_identical (183 to 212), check_manifest (225 to 239), compute_payload_pct (251 to 274) with check_payload (286 to 296), check_traceability (299 to 308) and check_agent_context (311 to 319), run_gate (322 to 339) and main (345 to 393). freeze_non.py: tree_sha256 (41 to 44), ignore (47) and main (52 to 145); freeze_non_parity_repair.py: the same three at 38 to 45, 48 and 53 to 142. build_phases.py is one script body (36 to 84) with gate_green (49 to 51). lint_hidden_scorers.py: lint_task_library (40 to 48), scanned_tasks (51) and main (56 to 102). lint_task_briefs.py: the agent-facing names (50), the retired prompts (57 to 60), the eight banned patterns with their labels (64 to 88), scan_text (91 to 99), scan_task_dir (102 to 112), scan_library (115 to 143), scan_prompts (146 to 156) and main (159 to 206). lint_task_literals.py: the stop words and patterns (60 to 69), ticket_words (72 to 79), literals_sent (82 to 153), uses_leniency (156), lint_task (161 to 197) and main (200 to 214). lint_retry_feedback.py: _norm (43), lint_task (47 to 108) and main (111 to 138). gen_retry_feedback.py: the three patterns (39 to 46), ticket_criteria (48 to 69), check_map (72 to 95), draft (98 to 122) and main (125 to 152). decontaminate_briefs.py: the twenty-one rules (28 to 94), decontaminate (97 to 113) and main (116 to 146). verify_task_absence.py: the tree locations (30 to 50), run_checks (53 to 86), the two classifiers (95 to 114) and main (117 to 175).
2.4 The library documents
The Project Library’s schema (project_manifest_schema.md, version 5b.project_manifest.v002) names a project’s required fields, its seam entries and its identifier formats; its README states what belongs in the library and the rule for maintenance experiments. The LAP Profile Library’s README lists the inventory, the status vocabulary and the agent-context index rule; LAP_LEVELS_GUIDE.md defines the eight rungs and the cumulative-nesting rule; measurement_policy.md fixes how a level’s payload is measured (characters divided by four against the source, the same model the instrumenter’s gate uses at instrument_profile.py 251 to 274); traceability_schema.md is the JSON Schema of traceability.json; lap_profile_index.csv records, per level and per built variant, the payload percentage and the build date. The Task Library’s README states the package anatomy and the registration discipline; task_manifest_schema.md carries two live versions, version 3 for the additive suite the campaign measures and version 2 for the retired removal-based suite whose thirty run sets must stay readable; seeded_defect_guide.md and task_scoring_rubric.md govern the retired suite’s defects and the pre-run factor scores; Project H/README.md names the seventeen live tasks and the two chained sequences.
3. The inputs
3.1 Command-line arguments
| Script | Parser | Arguments |
|---|---|---|
strip_lap.py | 160 to 174 | --src, --dst, --manifest (all required), --force |
materialize_feature_bases.py | 209 to 225 | --task (repeatable, required), --condition (default H-STR), --reference-condition (default the project’s LAP-L5P), --base-dir, --workdir, --allow-drift |
instrument_profile.py | 346 to 359 | --clean-code, --profile (one of the six registered), --out, --profile-lib (all required), --spec-dir, --model (default claude-opus-4-8), --dry-run |
build_phases.py | 36 to 40 | one positional, the project’s HASP Build folder |
lint_hidden_scorers.py | 57 to 68 | --project (default H), --pattern, --task-library, --verbose |
lint_task_briefs.py | 160 to 164 | --task-library (default 4. Task Library), --prompts, --quiet |
lint_task_literals.py | 201 to 204 | --task-library, --pattern (default *-M*) |
lint_retry_feedback.py | 112 to 114 | --library |
gen_retry_feedback.py | 126 to 130 | --library, --force, positional task identifiers |
decontaminate_briefs.py | 117 to 123 | positional library path, --apply, --preview DIR |
verify_task_absence.py | 118 to 124 | positional task folder, --witness, --witness-condition (default H-LAP-L5P), --quiet |
The two freeze scripts take no arguments; their paths are constants (freeze_non.py 34 to 36; freeze_non_parity_repair.py 29 to 32).
3.2 Environment variables
| Variable | Read at | What it decides |
|---|---|---|
FIVEB_LOCAL_BUILDS | verify_task_absence.py 43 to 47 | A local-disk copy of the three frozen builds, used when all three are present, because the repository’s volume is roughly 3,800 times slower for many small reads (comment, 34 to 42) |
FIVEB_PYTHON | verify_task_absence.py 48 to 49 | The interpreter that runs a task’s checks; default the experiment’s virtual environment |
FIVEB_VERIFY_RECORDS | verify_task_absence.py 73 to 75 | Where the checks’ per-level records land; default the temporary folder, because an earlier version wrote them beside the script inside a guarded subtree, which made verifying a task during a batch a way to halt a paid cell (comment, 59 to 72) |
FIVEB_CONDITION, FIVEB_WORKSPACE, FIVEB_HIDDEN_TIER_OUT, PYTHONPATH | set, not read: verify_task_absence.py 55 to 79 | The environment the scorer gives a task’s checks, reproduced exactly (chapter 08) |
HF_TOKEN, HF_HOME, LLM_HOME, SPARK_PORT | the spark/ scripts, chapter 03 | not this chapter’s |
No other script in this chapter reads an environment variable. The configuration files are read by path relative to the script folder (run_cell.py 68 and 1678 to 1679; score_cell.py 83; aggregate_metrics.py 109; preflight_batch.py 99).
3.3 Files read
| File | Read at | What it decides |
|---|---|---|
<variant>/lap_artifact_manifest.json | strip_lap.load_lap_files 146 to 156 and load_content_stub_files 126 to 143; materialize_feature_bases._strip_tree 85 to 87; instrument_profile.check_manifest 225 to 239 | Which files are documentation artefacts to delete or to count, and which behavioural content files are overwritten with a neutral placeholder rather than deleted |
3. LAP Profile Library/HASP Instrumentation/application_order.md and <profile>/artifact_rules.md, payload_budget.json, traceability_rules.md | instrument_profile.compose_prompt 89 to 92; load_payload_cap 277 to 283 | The instrumentation prompt and the level’s payload cap (uncapped for L5P) |
<task>/reference_feature.diff | materialize_feature_bases.materialize_task 153 to 157 | The feature-removal patch authored against the documented reference |
2. Project Library/*/variants/<reference> and <condition> | materialize_feature_bases.main 231 to 240 | The documented reference tree and the frozen target the derived diff must apply to |
<task>/task_manifest.json | lint_retry_feedback.main 118 to 127; preflight_batch.probe_task_registrations 374 to 383; analyze_coverage.py 152 to 157; run_cell.py 1826; score_cell.py 200 and 440; aggregate_metrics._load_task 373 to 386; task_sequences.py 137; gen_cell_narrative.py 161 | Whether the task is barred (valid_for_measured_runs), its scorer, its seeded defect, its chain predecessor |
<task>/TICKET.md, HOW_TO_VERIFY.md (and the retired task.md, visible_tests.md) | lint_task_briefs.scan_task_dir 102 to 112; lint_task_literals.ticket_words 72 to 79; gen_retry_feedback.draft 99 to 103; lint_retry_feedback.lint_task 65 to 68; decontaminate_briefs.main 127 to 130 | The agent-facing text that must not disclose the study, and the numbered criteria the retry notice may quote |
<task>/HIDDEN_CHECKS.md | gen_retry_feedback.check_map 72 to 95 at 104; lint_retry_feedback.lint_task 97 to 107 | The check-to-criterion map the specification declares |
<task>/hidden_checks/checks.py, checks_core.py | lint_task_literals.literals_sent 82 to 153; lint_retry_feedback.lint_task 90 to 96; verify_task_absence.main 127 to 130; lint_hidden_scorers through score_cell.describe_hidden_scorer 318 | Which literals the checks send, which check numbers they can fail, and whether an executable scorer exists |
<task>/retry_feedback.json | lint_retry_feedback.lint_task 50 to 56; gen_retry_feedback.main 138 to 141 (existence); run_cell.load_retry_feedback 271 to 288 (chapter 06) | The frozen map from a failed check to the ticket requirement it re-tests |
1. Harness/prompts/*.md | lint_task_briefs.scan_prompts 146 to 156; run_cell.py 1787 to 1788 and compose_artifact_hint 869 | The prompt of record and the hint template |
<build>/prompts/claude/phase-*.md, <build>/scripts/gate.sh | build_phases.py 41 and 49 to 51 | The corpus build’s phase prompts and its deterministic gate |
~/detached-jobs/5b-non-rebuild/workspace and iterate/state.json; ~/5b-non-repair-workspace | freeze_non.py 35, 53 and 75; freeze_non_parity_repair.py 30 and 59 | The control build’s working copy outside the repository, and the iteration loop’s spend record |
run_defaults.json | run_cell.py 1678 (max_passes, both budgets, 1686 to 1700); score_cell.py 83 (flaky_rerun_n); aggregate_metrics._load_defaults 107 to 117 (the budget ceiling and the pass-at-budget grid); preflight_batch.load_cost_model 97 to 105 | The pass ceiling, the budgets, the rerun count, the analysis grid, the cost estimate |
model_matrix.json | run_cell.py 1679 and 1685 | The default model class when no argument names one |
agent_handoff.schema.json | handoff_lib.py 11 (chapter 05) | The shape a handoff note is validated against |
hidden_judge.json | 4. Task Library/_hidden_lib/fiveb_judge.py 49 (chapter 08) | Whether the judging level runs and which model judges |
scoring_contract.json, reproducibility_schema.json | no program reads either: a search of every Python, shell and JavaScript file under 5. Experiment finds the first only in two comments of the scorer (score_cell.py 974 and 978) and the second nowhere | They are documentation of the rules the scorer and the runner carry in code (section 10) |
3.4 Network
Two scripts reach a model. instrument_profile.invoke_claude_cli (138 to 151) runs the agent tool with the prompt on standard input and saves standard output verbatim before anything parses it; build_phases.py (65 to 78) runs the same tool per phase under the subscription login with the user settings excluded (the comment at 18 to 25 explains why --bare would fail and why the setting sources are restricted). hidden_judge.json names the judge’s address for the local option (line 28) but is read by chapter 08’s library, not here. Every other script here reads and writes files only; the two freeze scripts also run pytest on the host (freeze_non.py 137 to 142; the parity repair through score_cell.compute_or_load_baseline_pass_set at 135 to 136).
4. The happy path in order
The order is the order in which the corpus and the task library are derived, then the checks that run before a batch, then how the configuration reaches a running cell.
4.1 The reference build, build_phases.py lines 36 to 84
The documented reference build of a project is produced by a loop over the phase prompts under the project’s HASP Build/prompts/claude/ (41), in sorted order. For each phase, a phase that already has an attempt transcript and a green gate is skipped (56 to 58), so the loop resumes; otherwise the agent tool is run with the phase prompt on standard input and its stream saved verbatim to generation_transcripts/<stem>.claude.attempt-<k>.jsonl, never overwriting an earlier attempt (59 to 78); then the host-run gate scripts/gate.sh <stem> must exit zero or the loop halts (82 to 83). The docstring states the structural role separation: the agent session implements, the deterministic gate tests, and an independent review pass reviews, because a script cannot grade itself (27 to 30). The gate re-run is the sole enforcement in a headless build, since no stop hook fires in a headless session (8 to 12). For Project H the frozen result is variants/H-LAP-L5P, which carries lap_artifact_manifest.json, traceability.json and AGENT_CONTEXT.md beside the code.
4.2 The strip, strip_lap.py lines 159 to 266
The stripped twin is derived, never rebuilt. main refuses a missing source or an existing destination without --force (178 to 183), reads the manifest’s artefact list and content stubs (185 to 186), copies the whole tree (189), deletes every listed artefact and records the ones that were missing (202 to 211), overwrites each behavioural content file with its neutral placeholder so the feature keeps working without the documentation text (213 to 219), and strips every Python file: docstrings first, because the syntax tree’s line spans are computed on the original source, then comments on the result (strip_python_source, 120 to 123). A docstring that is the sole statement of a function or class body is replaced by pass so the block stays valid (_docstring_spans, 66 to 70); a comment inside a string literal is never mistaken for a comment because the tokeniser decides (strip_comments, 97 to 101). Each stripped file is re-parsed to guarantee a broken file is never emitted (228); a file that will not parse is skipped and recorded, and the run exits 1 with a warning that the validity gate will fail (229 to 233, 258 to 264). The report strip_report.json records every deletion, stub, processed file with its hashes before and after, and skip (191 to 199, 245 to 256). The frozen H-STR carries that report at its root.
4.3 The instrumentation and its gate, instrument_profile.py lines 345 to 393
Each intermediate level is a fresh instrumentation of the clean code. main composes the prompt from the profile library (compose_prompt, 72 to 132): the application order, the level’s artefact rules, its payload budget and its traceability rules (89 to 92), an optional pointer to the requirements so anchors use the stable identifiers (94 to 100), the rule that AGENT_CONTEXT.md is required from L2S upward and forbidden at L1 (102 to 114), and the hard rules: exactly this level’s artefact kinds, all executable code byte-identical, nothing deleted, a root traceability.json, and a manifest listing every sidecar file created (116 to 131). The clean tree is copied to the output after every file is hashed (copy_clean_tree, 170 to 177), the agent tool is run with the transcript saved verbatim under an attempt number (138 to 160, 374 to 375), and a non-zero exit stops with the transcript’s path (376 to 378). Then run_gate (322 to 339) runs the five deterministic checks: every clean Python file, after stripping, must equal the clean file, and every non-Python file must be byte-identical (check_code_identical, 183 to 212, conservative by design because no deterministic stripper exists for other languages); the manifest must parse and every listed file exist (225 to 239); the payload, all sidecar characters plus the growth of code files, divided by the clean source’s characters, must be within the level’s cap unless the level is uncapped (251 to 296); traceability.json must parse (299 to 308); and the agent-context index must be present exactly when the level requires it (311 to 319). The report is written as instrumentation_gate_report.json (382 to 383); a red gate prints its failures and returns 1 (385 to 389). With --dry-run the prompt is written and no model is called (366 to 371).
4.4 The control freeze, freeze_non.py lines 52 to 145 and freeze_non_parity_repair.py lines 53 to 142
The control build is written independently from the specification and frozen by hand. freeze_non.py refuses to run twice or without a build (57 to 60), retires the earlier interim delivery by renaming it rather than deleting it (62 to 63, the North Star’s rule that prior states are preserved), copies the rebuilt tree without caches and dependency trees (65 to 66), hashes it under a convention it fixes in code because the earlier freeze’s hash could not be reproduced by any tried aggregation (7 to 20, 41 to 44), and writes variant_manifest.json with the digest, the builder, the six coverage audits and their reporting rule, the spend from the iteration loop’s state, the verification counts and the predecessor (74 to 134). The parity repair of 2026-08-22 (freeze_non_parity_repair.py) additionally verifies the outgoing frozen tree against its own manifest before retiring it, so the chain of custody is checked rather than assumed (64 to 70), copies the repaired workspace (75 to 76), extends the manifest with the repair record (82 to 128), and regenerates baseline_pass_set.json through the scorer’s own path (130 to 139). Both scripts are one-shot; nothing in the harness calls them.
4.5 The per-arm feature realisation, materialize_feature_bases.py lines 208 to 272 and 149 to 205
For the retired removal-based suite a feature task carried a patch that removes the feature from the documented reference, and a patch applies cleanly only to the tree it was written against. main copies the reference into a scratch folder and strips it (246 to 251), then checks that the stripped reference reproduces the frozen stripped variant file for file (verify_strip_matches_frozen, 112 to 124, at 253), refusing to generate anything against a drifted base unless told to allow drift (254 to 262); this is integrity check 1 of the docstring (14 to 22). materialize_task then applies the task’s removal patch to a fresh copy of the reference, refusing if it does not apply cleanly (159 to 165), strips the removed copy (167 to 168), diffs the two stripped trees to obtain the removal anchored on the stripped arm (170), refuses an empty diff because a feature whose removal strips to nothing would re-create run set 007’s vacuous staging (171 to 175), dry-runs the diff on a scratch copy of the frozen target (177 to 185), and writes reference_feature.<condition>.diff with a manifest carrying both hashes (187 to 205). The runner picks the diff up at staging (run_cell.py 2134 to 2158, chapter 06). The additive suite stages nothing and has no realisations; the three manifests under Project H/_archive/removal-based-suite-2026-08-22/ are the only ones in the library today.
4.6 The task registration and the retry map
A task under version 3 of the manifest schema is registered by hand: the ticket and the verification note are written first, the checks are drafted from the ticket alone and audited sentence by sentence, the checks are executed against the three untouched builds and must fail on all three, and only then is valid_for_measured_runs set true (task_manifest_schema.md, the paragraph beginning “What a task must have demonstrated”). verify_task_absence.py is that demonstration: it runs the task’s hidden_checks/checks.py against each of the three frozen trees in exactly the scorer’s environment (53 to 86), refuses a tree the checks pass on because the staging gate would then refuse every cell (140 to 143), and tells a genuine failure from a broken binding by the output’s shape, because both exit non-zero and the reference task in run set 039 was graded absent by a grader that never found the surface it was looking for (89 to 107, 147 to 157); with --witness it also demands a pass on an honest implementation (161 to 172). gen_retry_feedback.py then drafts retry_feedback.json from the ticket’s first numbered list (ticket_criteria, 48 to 69) and the check-to-criterion references in HIDDEN_CHECKS.md (check_map, 72 to 95), keeping an existing file unless forced because a person may have corrected it (138 to 141), and lint_retry_feedback.py verifies it (section 4.7).
4.7 The four lints and the decontaminator, run before a batch
lint_task_briefs.py scans the four agent-facing file names of every task folder found by either naming scheme, excluding the archive (115 to 143), and every prompt a current cell can select (146 to 156), for eight banned patterns ordered most specific first: the treatment’s name, a build or arm identifier, the study design, the hidden tier, an experiment concept, the word variant, a build-specific artefact, and the harness (64 to 88). It exits 1 on any hit and prints where designer notes belong (191 to 206); its docstring records the audit of 2026-08-15 that found seven of thirty-eight briefs describing the experiment to the agent (4 to 25). lint_hidden_scorers.py asks the scorer’s own classifier whether each task has an executable scorer (40 to 48; score_cell.describe_hidden_scorer 318) and exits 1 naming the tasks that do not, because their cells would be recorded as hidden_scorer_missing and never as agent failures (96 to 102). lint_task_literals.py reads every string literal the checks send from the syntax tree, never from text, so a word in a comment is never counted (82 to 153), separates names from values and discounts tolerant membership tests and docstrings (108 to 153), and reports a task whose checks send a value or name the ticket never presented as a literal and that does not go through the shared leniency helper (161 to 197); it deliberately does not guess which fix is right (37 to 39). lint_retry_feedback.py checks, for every registered task with a ticket, that the map exists and parses, that every mapped criterion is in the ticket verbatim and matches the ticket’s own numbered item, that no quoted criterion trips the brief lint, that every check number the executable grader can fail is mapped, and that the map agrees with the written specification (47 to 108); a barred task is reported but not counted (124 to 135). decontaminate_briefs.py is the one-time rewrite that removed the framing the audit found, with twenty-one ordered rules (28 to 94) and a preview mode; it keeps a .pre-1.11 copy beside each rewritten file (143 to 144) and must never be applied while a batch measures, because the library is a guarded subtree (3 to 6).
4.8 How the configuration reaches a cell
At run time the runner reads run_defaults.json and model_matrix.json once (run_cell.py 1678 to 1679): the pass ceiling is the argument or the default (1686), the budget is the argument, the dgx_claude figure, or the default (1694 to 1700), and the model is the argument, the Fabric alias, or the matrix’s default class (1681 to 1685). The prompt of record is the constant SINGLE_PROMPT (104), read at 1788 and copied into the run set’s prompts_snapshot/ (2222) so that every run set carries the exact text its cells ran under. The task manifest is read at 1826 for its seeded defect and its scorer. The scorer reads flaky_rerun_n (score_cell.py 83). The aggregator reads the budget ceiling and the grid (aggregate_metrics.py 107 to 117). The preflight reads the cost model to print a batch’s worst case before launch (preflight_batch.py 97 to 105) and each named task’s manifest to refuse a barred task (374 to 383). The coverage analyser lists the live tasks by the same flag (analyze_coverage.py 152 to 157). The handoff library reads the schema (handoff_lib.py 11) and the hidden library reads the judge’s settings (fiveb_judge.py 49).
5. The state machines of a task and of a variant
The flow model has no chapter 01 subgraph. The first diagram’s states are the values of status and valid_for_measured_runs in task_manifest.json under both schema versions, and its transitions are the documents and scripts that move a task between them. Proposed node identifiers for a later stage of the model: task.draft, task.registered, task.frozen, task.ready, task.barred.
stateDiagram-v2 state "version 2 (retired removal-based suite)" as V2 { state "draft" as D2 state "registered" as R2 state "frozen (valid_for_measured_runs true)" as F2 [*] --> D2 : task authored; factor scores provisional (task_manifest_schema.md, status lifecycle 1) D2 --> R2 : realism gate, adversarial review, registration manifest in HASP Task Design/manifests (lifecycle 2) R2 --> F2 : per-arm realizations with executed baseline_failure_check; materialize_feature_bases.py 149..205 writes reference_feature.<arm>.diff (lifecycle 3) } state "version 3 (additive suite)" as V3 { state "draft (valid_for_measured_runs false)" as D3 state "ready (valid_for_measured_runs true)" as RD3 state "ready but barred (valid_for_measured_runs false)" as B3 [*] --> D3 : TICKET.md, HOW_TO_VERIFY.md, HIDDEN_CHECKS.md, checks_core.py drafted from the ticket alone D3 --> RD3 : verify_task_absence.py fails on all three untouched builds (140..157) and passes on a witness (161..172); gen_retry_feedback.py 98..122 writes the map; the four lints exit 0 RD3 --> B3 : a problem under investigation; the flag wins over the status (schema, "The two flags must disagree in exactly one direction") B3 --> RD3 : problem closed } F2 --> [*] : archived under Project H/_archive on 2026-08-22; manifests still read RD3 --> Batch : preflight_batch.probe_task_registrations 374..383 admits; analyze_coverage 152..157 counts; run_cell 1826 reads the manifest but never the flag B3 --> Refused : preflight 383 refuses the batch; lint_retry_feedback 124..135 reports but does not count
The second diagram’s states are the folders a product build passes through, and its transitions are the scripts of section 4 with the record each leaves. Proposed identifiers: corpus.reference_built, corpus.reference_frozen, corpus.stripped, corpus.instrumented, corpus.control_frozen, corpus.realisation_written.
stateDiagram-v2 [*] --> Reference_building : build_phases.py 54..84, one phase per prompt, gate.sh must exit 0 (82..83) Reference_building --> Reference_frozen : all phases green; variants/H-LAP-L5P with lap_artifact_manifest.json, traceability.json, AGENT_CONTEXT.md Reference_frozen --> Stripped : strip_lap.py main 159..266; strip_report.json 255..256; variants/H-STR Reference_frozen --> Instrumenting : instrument_profile.py copy_clean_tree 170 from the stripped tree; claude run 375 Instrumenting --> Instrumented : run_gate 322..339 green; instrumentation_gate_report.json 382..383; variants/H-LAP-L1 .. L4P Instrumenting --> Red_gate : any of the five checks fails; report written; exit 1 (385..389) Red_gate --> Instrumenting : re-run; transcripts keep their attempt numbers (154..160) [*] --> Control_building : the iterate-to-spec loop outside the repository (freeze_non.py 35, 90..94) Control_building --> Control_frozen : freeze_non.py 52..145; variant_manifest.json 133..134; the earlier delivery retired beside it 62..63 Control_frozen --> Control_frozen : freeze_non_parity_repair.py 53..142 verifies the outgoing hash 64..70, retires it, refreezes, regenerates baseline_pass_set.json 130..139 Stripped --> Realisation_written : materialize_feature_bases.py 208..272 for a removal-based feature task; reference_feature.H-STR.diff and its manifest 187..204 Stripped --> Strip_drift : verify_strip_matches_frozen 112..124 finds a mismatch; refused unless --allow-drift 254..262
6. The sequence of the instrumentation gate
The participants are the operator, the instrumenter, the profile library, the agent tool, and the files. This is the diagram plan’s third element for the chapter; the flow model has no sequence for it.
sequenceDiagram participant O as operator participant I as instrument_profile.main 345 participant L as 3. LAP Profile Library participant A as claude -p (invoke_claude_cli 138) participant X as variant output folder O->>I: --clean-code H-STR --profile L2S --out H-LAP-L2S --profile-lib ... [--spec-dir] [--dry-run] I->>L: application_order.md, artifact_rules.md, payload_budget.json, traceability_rules.md (compose_prompt 89..92) I->>I: agent-context rule required (L2S..L5P) or forbidden (L1) 102..114; hard rules 116..131 I->>X: copy the clean tree after hashing every file (copy_clean_tree 170..177) alt --dry-run I->>X: _instrumentation_prompt.txt 368; return 0 else I->>A: prompt on stdin, cwd = out, --model 143..146 A-->>X: generation_transcripts/instrument-<profile>.claude.attempt-<k>.jsonl 147 (and .stderr.txt 149) A-->>X: sidecar documentation files, inert comments, traceability.json, lap_artifact_manifest.json, AGENT_CONTEXT.md A-->>I: exit status; non-zero stops with the transcript path 376..378 I->>I: check_code_identical 183..212 (strip(out) == clean per .py; bytes equal otherwise) I->>X: read lap_artifact_manifest.json (check_manifest 225..239) I->>I: compute_payload_pct 251..274 against load_payload_cap 277..283 I->>X: read traceability.json (check_traceability 299..308) I->>I: check_agent_context 311..319 I->>X: instrumentation_gate_report.json 382..383 alt failures empty I-->>O: gate GREEN with the payload percentage 390..393; return 0 else I-->>O: RED GATE with each failure on stderr 385..389; return 1 end end
7. The records, with their writers and readers
Every file this chapter’s scripts write, and the six configuration files and six prompts as records whose writer is the operator’s hand. Readers are named by file and line where the reading is one call and by file where it is spread. Every record here is single agent only.
| Record | Writer | Fields or content | Readers |
|---|---|---|---|
1. Harness/config/run_defaults.json | the operator, by hand; the basis of each figure is written beside it | section 2.1 | run_cell.py 1678 to 1700; score_cell.py 83; aggregate_metrics._load_defaults 107 to 117; preflight_batch.load_cost_model 97 to 105 |
1. Harness/config/model_matrix.json | the operator | section 2.1 | run_cell.py 1679 and 1685 (default_model_class only; the pinned identifier is recorded from transcripts, not read from here) |
1. Harness/config/scoring_contract.json | the operator | section 2.1 | none by code; the seven validity reasons are carried in score_cell.collect_invalid_reasons 976 (chapter 08) and the file is cited in its comments (974, 978) |
1. Harness/config/reproducibility_schema.json | the operator | section 2.1 | none by code (section 3.3, last row) |
1. Harness/config/agent_handoff.schema.json | the operator | section 2.1 | handoff_lib.py 11 (chapter 05), through which run_cell.capture_handoff 1156 to 1193 validates a note without affecting the verdict |
1. Harness/config/hidden_judge.json | the operator | section 2.1 | 4. Task Library/_hidden_lib/fiveb_judge.py 49 (chapter 08) |
1. Harness/prompts/maintenance.single_agent.v006.md and the five others | the operator; v006 ratified as amendment A12, Decision Sheet item 55, 2026-09-03 (run_cell.py 92 to 103) | section 2.2 | run_cell.py 1786 to 1788 (the prompt of record), compose_artifact_hint 869 to 922 (the hint), 2222 (the snapshot into <run set>/prompts_snapshot/); lint_task_briefs.scan_prompts 146 to 156 |
<variant>/strip_report.json | strip_lap.main 255 to 256 | source, destination, deleted_lap_files, missing_lap_files, content_stubbed_files, python_files_processed (each with file, before, after, changed), skipped_for_syntax_error, counts | materialize_feature_bases._tree_hashes 104 excludes it from the comparison; readers of the frozen H-STR |
<variant>/lap_artifact_manifest.json | the instrumenting agent, under the hard rule at instrument_profile.py 128 to 130; for the reference, the build’s own instrumentation | lap_files (or the older artifacts), trace_anchor_sites, probe_comment_sites (legacy), content_stub_files, payload_metrics | strip_lap.load_lap_files 146; instrument_profile.check_manifest 225; materialize_feature_bases._strip_tree 85; ledger_lib.load_lap_manifest 212 (chapter 09); run_cell.py 2209 and compose_artifact_hint 869 (chapter 06); aggregate_metrics.load_snapshot 152 (chapter 10) |
<variant>/instrumentation_gate_report.json | instrument_profile.main 382 to 383 | profile, ok, failures, payload_pct_chars_div4, payload_cap, lap_files, transcript | the operator; lap_profile_index.csv records the payload figure by hand |
<variant>/_instrumentation_prompt.txt | instrument_profile.main 368 (dry run only) | the composed prompt | the operator |
<project>/HASP Build/generation_transcripts/<stem>.claude.attempt-<k>.jsonl and instrument-<profile>.claude.attempt-<k>.jsonl (with .stderr.txt) | build_phases.py 62 to 78; instrument_profile.invoke_claude_cli 147 to 150 | The agent tool’s stream, verbatim, one file per attempt, never overwritten | the generation ledger of 6. Metrics/generation_ledger_spec.md section 6 (no script in the harness builds it today, chapter 09 section 14.3) |
variants/H-NON/variant_manifest.json | freeze_non.py 133 to 134; extended by freeze_non_parity_repair.py 127 to 128 | variant, rebuilt, frozen_utc, tree_sha256, tree_sha256_convention, file_count, builder_model, builder_session, coverage (six audits and the reporting rule), spend, verification, predecessor, predecessor_note, metadata_files_outside_tree_hash, and after the repair repaired, repair (date, defects closed, source, builder, acceptance gate, report, files changed) | freeze_non_parity_repair.py 62 to 70 (the outgoing hash check); the plan’s Appendix L; readers of the control build |
variants/<variant>/baseline_pass_set.json | freeze_non_parity_repair.py 130 to 139 through score_cell.compute_or_load_baseline_pass_set 765 (chapter 08 owns the writer) | pass_ids | score_cell.py (the full-suite regression gate, scoring_contract.json 44 to 54) |
<task>/reference_feature.<condition>.diff and reference_feature.<condition>.manifest.json | materialize_feature_bases.materialize_task 187 to 204 | the unified diff; schema_version 5b.feature_realization.v001, task_id, condition, generated, method, source_removal_diff, source_removal_sha256, generated_diff, generated_diff_sha256, dry_run_clean_on_frozen_variant | run_cell.py 2134 to 2158 (chapter 06); the archived suite only |
<task>/retry_feedback.json | gen_retry_feedback.main 146 to 147, from draft 98 to 122; corrected by hand only to fix a mapping (the file’s own _note) | schema_version 5b.retry_feedback.v001, task_id, _note, checks (check number to criterion), criteria (criterion number to the ticket’s sentence, verbatim) | lint_retry_feedback.lint_task 50 to 108; run_cell.load_retry_feedback 271 to 288 and unmet_criterion 291 to 316 (chapter 06) |
<task>/task.md.pre-1.11, visible_tests.md.pre-1.11 | decontaminate_briefs.main 143 (with --apply) | the brief as it was before the rewrite | the operator, for the diff |
<FIVEB_VERIFY_RECORDS>/<task>.<condition>.json | a task’s checks, through the path verify_task_absence.run_checks sets in FIVEB_HIDDEN_TIER_OUT (77 to 78); the scorer sets the same variable at run time (chapter 08) | which level of evidence answered, per check | the operator |
<task>/task_manifest.json | the task author, by hand, to task_manifest_schema.md | version 3: schema_version, task_id, project_id, task_type, difficulty_band, seeded_defect.present, hidden_scorer (type, script), documentation_obligation, status, valid_for_measured_runs, and for a chained step depends_on with depends_on_note | section 3.3 |
The lineage from these records to the programs that consume them is drawn below. Every arrow is a reader relationship from the table above or from section 3.3.
flowchart LR subgraph config [1. Harness/config, written by the operator] RD[run_defaults.json] MM[model_matrix.json] SC[scoring_contract.json<br/>no reader in code] RS[reproducibility_schema.json<br/>no reader in code] HS[agent_handoff.schema.json] HJ[hidden_judge.json] end subgraph prompts [1. Harness/prompts] P6[maintenance.single_agent.v006.md] PH[artifact_hint.v001.md] end subgraph corpus [2. Project Library variants] L5[H-LAP-L5P<br/>build_phases.py 54..84] STR[H-STR<br/>strip_lap.py 159..266; strip_report.json] LN[H-LAP-L1 .. L4P<br/>instrument_profile.py; instrumentation_gate_report.json 382] NON[H-NON<br/>freeze_non.py 133; parity repair 127; variant_manifest.json] LM[lap_artifact_manifest.json per documented variant] end subgraph tasks [4. Task Library/Project H/<task>] TM[task_manifest.json] TK[TICKET.md, HOW_TO_VERIFY.md] HC[HIDDEN_CHECKS.md, hidden_checks/checks.py, checks_core.py] RF[retry_feedback.json<br/>gen_retry_feedback.py 146] RFD[reference_feature.H-STR.diff, archive only<br/>materialize_feature_bases.py 187..204] end L5 --> STR STR --> LN L5 --> RFD STR --> RFD LINT1[lint_task_briefs.py 159] LINT2[lint_hidden_scorers.py 56] LINT3[lint_task_literals.py 200] LINT4[lint_retry_feedback.py 111] VER[verify_task_absence.py 117] TK --> LINT1 P6 --> LINT1 PH --> LINT1 HC --> LINT2 TK --> LINT3 HC --> LINT3 RF --> LINT4 TK --> LINT4 HC --> LINT4 HC --> VER L5 --> VER STR --> VER NON --> VER RC[run_cell.py 1678, 1679, 1788, 1826, 2089, 2134, 2209, 2222, 271] RD --> RC MM --> RC P6 --> RC PH --> RC TM --> RC TK --> RC RF --> RC RFD --> RC LM --> RC L5 --> RC STR --> RC LN --> RC NON --> RC SCR[score_cell.py 83, 200, 318, 440] RD --> SCR TM --> SCR HC --> SCR HJ --> JUDGE[_hidden_lib/fiveb_judge.py 49] HS --> HL[handoff_lib.py 11] PF[preflight_batch.py 97..105, 374..383] RD --> PF TM --> PF AG[aggregate_metrics.py 107..117, 152, 373..386] RD --> AG TM --> AG LM --> AG AC[analyze_coverage.py 152..157] TM --> AC LL[ledger_lib.py 47..52, 212] LM --> LL STRIP[strip_lap.strip_python_source 120] --> LL
8. The loops and the waits
No script here waits on anything external except the two that run the agent tool, which block until the tool exits with no timeout (instrument_profile.py 143 to 146; build_phases.py 65 to 78), and the absence verifier, which gives each check run 600 seconds and records a timeout as exit status 124 (verify_task_absence.py 50, 80 to 85). The loops are over files: every listed artefact and every Python file in the strip (strip_lap.py 202 to 243); every phase prompt in the build (build_phases.py 54 to 83), which resumes because a phase with a transcript and a green gate is skipped; every task named on the command line in the materialiser (materialize_feature_bases.py 264 to 271); every task folder found by either brief name and every prompt in the brief lint (115 to 156); every task folder with checks in the literal lint (206 to 209); every manifest with a ticket in the retry lint (118 to 135); every folder with a checks specification in the generator (132 to 148); every agent-facing file in the decontaminator (127 to 144); and the three frozen trees in the verifier (135 to 157). The attempt numbering of a transcript is a loop that finds the first unused number (instrument_profile.py 157 to 160; build_phases.py 59 to 61). No script retries anything.
9. The guards and refusals
| Guard | Where | What it refuses or records |
|---|---|---|
| The strip never emits a broken file | strip_lap.py 225 to 233 | A stripped file that will not re-parse is skipped and recorded; the run exits 1 |
| The strip never overwrites silently | 180 to 183 | An existing destination is refused without --force |
| A malformed manifest is refused | load_content_stub_files 136 to 142; load_lap_files 153 to 156 | An unrecognised schema or a stub map that is not a string-to-string object |
| A comment inside a string is never a comment | strip_comments 97 to 101 | The tokeniser decides, never a pattern |
A sole docstring becomes pass | _docstring_spans 66 to 70 | So a function or class body stays valid |
| The stripped reference must reproduce the frozen twin | materialize_feature_bases.verify_strip_matches_frozen 112 to 124, at 253 to 262 | A file only on one side or with different content refuses generation unless --allow-drift |
| The removal patch must apply cleanly | 163 to 165 | Otherwise the task is refused |
| The derived diff must not be empty | 171 to 175 | A feature that strips to nothing would hand the agent a tree that already satisfies the task |
| The derived diff must dry-run clean on the frozen target | 177 to 185 | Otherwise the stripped reference has drifted |
| A generated diff is never hand-edited | docstring 27 | The script is re-run instead |
| Executable code stays byte-identical across levels | instrument_profile.check_code_identical 183 to 212 | A stripped mismatch in a Python file, any change in another file, or a deleted file fails the gate |
| The manifest must list files that exist | check_manifest 225 to 239 | A missing manifest, a parse error, or a listed file that is absent |
| The payload stays within the level’s cap | check_payload 286 to 296 | Over the cap fails; L5P is uncapped |
traceability.json must parse | check_traceability 299 to 308 | A missing or unparsable file |
| The agent-context index is present exactly when required | check_agent_context 311 to 319 | Present at L1, or absent from L2S upward, fails |
| A red gate never passes silently | main 385 to 389; docstring 20 to 21 | The failures are printed and the exit is 1 |
| A model exit that is not zero stops the instrumentation | 376 to 378 | With the transcript path in the message |
| The build’s gate is the sole enforcement | build_phases.py 82 to 83; 8 to 12 | A red gate halts; a green gate advances; a phase is never re-run once green |
| Transcripts are never overwritten | build_phases.py 59 to 62; instrument_profile.py 154 to 160 | Attempt numbers only ever increase |
| A freeze never runs twice or on nothing | freeze_non.py 57 to 60; freeze_non_parity_repair.py 57 to 60 | The retired name existing, or the source missing, refuses |
| The outgoing frozen tree must match its manifest | freeze_non_parity_repair.py 64 to 70 | Otherwise nothing is frozen on top of it |
| A brief may not disclose the study | lint_task_briefs.BANNED 64 to 88; run_cell.py 2089 | Any of eight patterns in an agent-facing file or a selectable prompt; the runner refuses the cell before any spend |
| The retired prompts are not launchable and not linted | lint_task_briefs.py 52 to 60 | They are kept for old records only |
| Archived tasks are not scanned here but are at run time | scan_library 124 to 131 | The runner’s per-cell scan remains universal |
| A task must have an executable scorer | lint_hidden_scorers.lint_task_library 40 to 48 | Otherwise its cells would be hidden_scorer_missing, an instrument error |
| A check may not demand a spelling the ticket did not quote | lint_task_literals.lint_task 161 to 197 | Unless the check goes through the leniency helper; the lint refuses to guess which fix is right |
| Literals are read from the syntax tree | literals_sent 91 to 99 | A word in a comment or docstring is never counted |
| The retry map may only repeat the ticket | lint_retry_feedback.lint_task 76 to 87 | A paraphrase, a missing criterion, or a criterion that trips the brief lint |
| Every check the grader can fail is mapped | 89 to 96 | A _fail(N, ...) in checks_core.py with no entry |
| The map agrees with the written specification | 97 to 107 | A check number or criterion that differs from HIDDEN_CHECKS.md |
| A barred task is reported, not counted | 124 to 135 | valid_for_measured_runs false exempts it from the exit status |
| A drafted map is never overwritten silently | gen_retry_feedback.main 138 to 141 | Unless --force |
The decontaminator writes only with --apply and keeps a copy | decontaminate_briefs.main 138 to 144 | Preview or dry run otherwise; never during a batch |
| A passing untouched build means a vacuous task | verify_task_absence.main 140 to 143 | The staging gate would refuse every cell |
| A crash is not an absence demonstration | 89 to 107, 150 to 157 | A traceback or binding error is BROKEN BINDING, not fails as required |
| The verifier’s records never land in a guarded subtree | 59 to 78 | The temporary folder, or FIVEB_VERIFY_RECORDS |
| The flag wins over the status | task_manifest_schema.md; preflight_batch.py 383; analyze_coverage.py 157 | A task barred by flag is refused whatever its status says |
10. Every unhappy path, in four parts
Each row states what the step is supposed to do and why it works that way, what goes wrong with the trigger and the code path, what is written with the status and exit, and what it costs downstream.
| Supposed to do, and why it works that way | What goes wrong: trigger and code path | What is written; status and exit | What it costs downstream |
|---|---|---|---|
| Strip every Python file to its executable code, re-parsing each result so a broken file is never emitted | A file the parser accepts before stripping fails to parse after; strip_lap.py 225 to 233 | The file is left unstripped; strip_report.json lists it under skipped_for_syntax_error; exit 1 with a warning | The stripped variant carries one file with its comments; the instrumenter’s gate would then report a stripped mismatch for that file on every level; the strip must be fixed and re-run |
| Delete every documentation artefact the manifest lists | A listed file is absent from the source; 202 to 211 | Recorded under missing_lap_files; the run continues, exit 0 | Nothing, unless the manifest was meant to list a file that exists under another name, which the report makes visible |
| Derive a removal patch that applies to the stripped arm by construction | The stripped reference no longer reproduces the frozen stripped variant, because one of them was edited since the freeze; verify_strip_matches_frozen 112 to 124 at 253 to 262 | Each mismatch printed as STRIP-DRIFT; SystemExit, nothing written | No realisation for the task; a cell staged for it would be refused as missing_feature_realization (chapter 06). With --allow-drift a diff is generated against a drifted base and may fail the dry run at 181 to 185 instead |
| Refuse a removal that strips to nothing | The feature lives entirely in documentation, so the two stripped trees are identical; 171 to 175 | SystemExit naming the task; nothing written | The task cannot be measured on the stripped arm; this is the run set 007 defect the docstring names (22) |
| Instrument the clean code without changing it | The model changed executable code, deleted a file, or changed a non-Python file; check_code_identical 183 to 212 | instrumentation_gate_report.json with ok false and the failures; exit 1 | The level is not usable; the operator re-runs, and the next transcript takes the next attempt number. Nothing in the harness reads a red report, so a red variant left in the library would be indistinguishable from a green one to the runner, which reads only the tree; the report and lap_profile_index.csv are the only records |
| Keep the level’s payload within its cap | The model wrote more documentation than the cap; check_payload 286 to 296 | as above | as above; the figure is a characters-divided-by-four estimate, the policy’s own model |
| Run the model to instrument a level | The agent tool exits non-zero; 376 to 378 | The transcript and standard error are saved (147 to 150); SystemExit naming the transcript; no gate, no report | The output folder holds a partial tree; a re-run copies the clean tree afresh (174 to 176) |
| Build the reference phase by phase with a deterministic gate | The gate exits non-zero after a phase; build_phases.py 82 to 83 | The attempt transcript is kept; SystemExit naming the phase | The build halts; a re-run skips green phases and retries the red one under a new attempt number |
| Freeze the control once, preserving the earlier state | The retired name already exists, or the source build is absent; freeze_non.py 57 to 60 | SystemExit; nothing moved | The freeze must be reconciled by hand; the earlier delivery is never overwritten |
| Verify the outgoing frozen tree before freezing on top of it | The tree’s hash differs from its manifest; freeze_non_parity_repair.py 64 to 70 | SystemExit naming both hashes | Someone edited a frozen variant; the plan’s read-only rule for variants was broken and must be investigated before any further freeze |
| Keep every agent-facing file free of the study | A brief or a selectable prompt matches a banned pattern; lint_task_briefs.main 191 to 206; run_cell.py 2089 to 2092 at run time | The hits printed with file, line and label; exit 1. At run time the cell is refused as contaminated_task_brief:<file>:<line>:<label> with a NOT-RUN line and no spend (chapter 06) | A batch over the task cannot run until the brief is rewritten; the lint’s docstring records the seven briefs of 2026-08-15 that had already run in thirty run sets |
| Prove every task has an executable scorer | A task has neither a registered asset nor a probe block; lint_hidden_scorers.main 96 to 102 | The offenders on standard error; exit 1; exit 2 when the library folder is missing | A cell on such a task would be recorded hidden_scorer_missing and invalid; run set 007’s false 0 of 5 is the incident (scoring_contract.json 28) |
| Refuse a check that demands a spelling the ticket never demanded | The checks send a bare word the ticket never quoted and do not use the leniency helper; lint_task_literals.lint_task 177 to 196 | A VALUE SPELLING or FIELD NAME finding per task; exit 1 | Left unfixed, correct work is failed for a spelling habit; the docstring records sixteen of nineteen failures on H-M021 across run sets 035 to 039 and about 125 million tokens of agent time (13 to 25) |
| Keep the retry map complete and honest | The map is missing, unreadable, maps nothing, quotes a paraphrase, omits a check the grader can fail, trips the brief lint, or disagrees with the specification; lint_retry_feedback.lint_task 47 to 108 | The problems per task; exit 1 unless the task is barred | A cell whose hidden failure has no mapping falls back to the generic notice and records that it did (run_cell.py 291 to 316), making its attempts incomparable with cells that received the informative notice (docstring, 8 to 11) |
| Draft the map from the ticket’s numbered list | The ticket’s list does not start at 1 or is not the first numbered run, or the specification’s criterion references do not match the three patterns; ticket_criteria 48 to 69; check_map 72 to 95 | A map with missing checks or criteria, written unless the file exists | The lint above catches it before any cell runs |
| Demonstrate that a task’s checks fail on every untouched build | A tree is missing, a check run times out, the checks pass on an untouched build, or the failure is a crash rather than a named check; verify_task_absence.main 135 to 157 | The per-build line and verdict: NOT READY; exit 1; exit 2 when the task has no checks.py | The task must not be set valid_for_measured_runs; a passing untouched build means the staging gate would refuse every cell; a crash means the grader cannot reach the product, which is what graded correct work as absent in run set 039 (89 to 94) |
| Keep the read-only configuration consistent with the code that carries the same rules | scoring_contract.json lists seven validity reasons and reproducibility_schema.json lists the manifest fields, and no program reads either; the scorer’s collect_invalid_reasons (score_cell.py 976, chapter 08) and the runner’s manifests carry the rules in code | Nothing; the files change only by hand | The two files can drift from the code without any check noticing. metric_definitions.md lines 131 to 139 already print one validity rule retired on 2026-09-03 (operations/results-integrity/B2-count-provenance.md line 7), which is the same class of drift. The cost is a reader who trusts the file over the code |
| Give a run set the prompt its cells ran under | The prompt of record changes between two cells of one run set | Each cell’s prompt_file in metrics.json names the file; the run manifest lists every version seen (run_cell.py 3168 to 3173) | A run set that mixes versions is visible in its manifest; the snapshot at 2222 keeps only the last text per name |
11. The metrics this chapter produces or feeds
This chapter produces no metric; it fixes the constants others compute against and the corpus they measure. From run_defaults.json: max_passes bounds passes_run and passes_to_success; token_budget_per_cell and its dgx_claude counterpart bound cumulative_spend and set budget_exhausted (run_cell.py 2468); pass_at_budget_grid and the ceiling are the axes of pass_at_budget.json and group_analysis.json (aggregate_metrics.py 107 to 117, 955 to 997); flaky_rerun_n sets the flakiness rerun (score_cell.py 83); primary_metric names tokens_to_success. From model_matrix.json: the model_class column of every table when no argument names a model. From the task manifest: task_type, difficulty_band, documentation_obligation and acceptable_edit_set reach the matrix through aggregate_metrics._load_task (373 to 386) and compute_cell (task_type, difficulty_band at 641 to 642; the obligation and the cone at 410 to 411 and 429). From the profile library: lap_profile_id and the payload figures of lap_profile_index.csv and measurement_policy.md are the dose axis of the analysis (chapter 12), and the per-variant payload_metrics in lap_artifact_manifest.json are the denominator of dose_consumed_fraction (aggregate_metrics.payload_tokens 165 and 484 to 485). From the prompt: prompt_variant, artifact_hint_applied and retry_feedback_mode are analysis dimensions (chapter 06, section 11). From the retry map: retry_feedback_available and the per-pass retry_feedback record (chapter 06).
12. The tests that exercise the mechanism
test_instrument_profile.py (seventeen tests, lines 85 to 249) covers the prompt’s contents, the agent-context rule in both directions, the five gate checks each in its failing direction (executable code changed, a file deleted, a non-Python file changed, over the cap, uncapped L5P, missing manifest, manifest listing a missing file, bad traceability), the L1 and L2S agent-context cases, that a dry run makes no model call, and that main runs the gate through the seam. test_materialize_feature_bases.py (three tests, 96 to 136) covers a generated diff that removes the feature, the refusal on strip drift, and the refusal of an empty diff. test_lint_hidden_scorers.py (five tests, 33 to 106) covers a clean library, a task without a runnable probe, the pattern and missing-library cases, the reuse of the scorer’s classification, and the real Project H library’s state. test_prompt_variant_and_retry_feedback.py (thirty tests, 60 to 501) covers the hint’s composition and neutrality, the variant’s recording and forwarding, the criterion notice, and seven tests of the retry-map lint at 389 to 437 (a clean map, a missing map, a paraphrase, an unmapped failable check, disagreement with the specification, a criterion naming the study, and that every live task carries a complete map). test_isolation_contract.py covers the brief lint on a contaminated and a clean brief (69 and 88) and that the composed task package never contains hidden-check source (50). test_handoff_lib.py (seven tests, 61 to 114) covers the handoff schema’s required members and the validator’s behaviour on bad input. test_run_cell.py and test_score_cell.py read run_defaults.json through the runner and the scorer.
No test covers strip_lap.py directly (its behaviour is exercised through the materialiser’s tests and the ledger’s comment-only rule in test_classify_event.py, chapter 09), build_phases.py, the two freeze scripts, lint_task_literals.py, gen_retry_feedback.py’s parsers on their own, decontaminate_briefs.py, or verify_task_absence.py; and nothing checks scoring_contract.json or reproducibility_schema.json against the code that carries the same rules.
13. The dated incidents that shaped the code
The instrumenter is dated 2026-07-17 and quotes the design lock from the plan’s chapter 4 line 79 and the profile library’s README line 82 (instrument_profile.py 2 to 9); the build loop and the choice of subscription login over the bare mode are dated the same day (build_phases.py 2, 18). The materialiser is the amendment of 2026-08-08 after run set 007’s three feature cells were found staged with the feature already built (materialize_feature_bases.py 4 to 12, 21 to 22); its treatment of the coverage package name records the 2026-08-15 change to deciding artefact folders by contents (56 to 65). The control freeze of 2026-08-18 replaced the interim delivery of 2026-08-11 and fixed the tree-hash convention in code because the earlier hash could not be reproduced (freeze_non.py 3 to 20); the parity repair of 2026-08-22 closed four control-only defects found by the behavioural parity audit and added the outgoing-hash check (freeze_non_parity_repair.py 3 to 15). The brief lint records the audit of 2026-08-15 that found seven of thirty-eight briefs disclosing the study, with the exact briefs named (lint_task_briefs.py 4 to 25), and the discovery of 2026-08-27 that scanning by task.md alone had left the seventeen additive packages unscanned while the lint reported the library clean (118 to 122). The hidden-scorer lint and the hidden_scorer_missing reason date from 2026-08-08 after run set 007’s false 0 of 5 (lint_hidden_scorers.py 4 to 9; scoring_contract.json 24 to 29). The literal lint is the ruling of 2026-08-29 after run set 039, with the H-M021 arithmetic (lint_task_literals.py 2 to 25). The retry map is amendment A9 of 2026-09-01 (gen_retry_feedback.py 4 to 9). The shell_unavailable reason is Decision Sheet item 52 of 2026-09-03, with the 134 affected cells listed (scoring_contract.json 30 to 43). The one prompt for every build is v005 of 2026-09-02 and v006 of 2026-09-03 (amendment A12, Decision Sheet item 55), and the runner’s comment explains why a per-build prompt must not return (run_cell.py 81 to 103); the artifact hint is amendment A8 of 2026-09-01, after 71 of 71 survey cells on the documented builds had opened no documentation file (110 to 120). The dgx_claude budget is the ruling of 2026-08-28 executing Appendix N ruling N.4 (run_defaults.json 7). The judge was built under the ruling of 2026-08-29 (hidden_judge.json 9, 50). The measured model class was pinned on 2026-08-09 from the containerised run sets’ transcripts (model_matrix.json 9). The absence verifier’s record location was moved out of the guarded subtree after an earlier version made verifying a task during a batch a way to destroy a paid cell (verify_task_absence.py 59 to 72), and its crash classifier dates from run set 039’s reference task (89 to 94).
14. The weakest claim, what was not checked, and the token line
The weakest claim
The weakest claim is the last row of section 3.3 and the corresponding row of section 10, that no program reads scoring_contract.json or reproducibility_schema.json. It rests on a search for the two file names across every Python, shell and JavaScript file under 5. Experiment; a reader that builds the path from parts, or a script outside that folder, would not have been found. The second weakest is the reading of the version 3 task lifecycle in section 5: the schema describes status as free text with three values in use, and the diagram draws them as states, which is a reading of the document and not of any code that enforces them.
14.3 What was not checked
Whether every variant folder under 2. Project Library/Project H .../variants/ carries an instrumentation_gate_report.json with ok true; only the folder listings were read. Whether the eight rungs in lap_profile_index.csv match the six profiles the instrumenter registers (PROFILE_DIRS names L1, L2S, L2T, L3B, L4P and L5P; the index also lists L1H and L2C, whose folders exist in the library but which the instrumenter cannot select), which may mean two rungs were instrumented by another path. The HASP Instrumentation/ folder’s own gates and manifests were not read. The library guide documents were read by heading only. The generation transcripts’ ledger (generation_ledger_spec.md) has no builder in the harness, as chapter 09 records.
14.4 Token line
The session that wrote this chapter had consumed 628,458 tokens of context by the time writing began, measured as the difference of the remaining-token counter (15,000,000 at the start of the session, 14,371,542 at the last reading before the file was written); of that, 455,732 had been spent before chapter 03 was written, and the remainder on this chapter’s materials: the six configuration files and six prompts whole, the thirteen scripts whole or by inventory, the task manifest schema’s lifecycle sections, the library documents’ headings, the readers of every configuration file, and the tests. No harness script was run and no model tokens were spent by the harness.