The source for the chapter is the harness code at the verified commit named above, the chapter 08 subgraph in 5. Experiment/11. Detailed Design/flow-model/flow_model.v001.json, and the owning-file digests. The diagrams below distinguish the flow model from the hand-authored structural diagrams. The state, sequence and record-lineage diagrams render the chapter 08 model subgraph where stated, while the ladder and the supporting structural relationships are hand-authored from the cited code elements.

1. The purpose and the position in the life of the experiment

A cell is one measured attempt by an AI coding agent to complete one task under one experimental condition. Scoring is the work done after that attempt: it examines the workspace the agent left, repeats the registered visible test, compares the workspace with the project’s frozen full suite, runs checks kept outside the agent’s view, and records whether the result can be counted. A hidden check is an automated assertion kept in the task library rather than shown to the agent. A verdict is the recorded conclusion for the cell. The cell runner invokes score_cell.py after the run, and the post-experiment rescoring program can invoke the hidden checks again against a preserved tree without running the agent again.

The mechanism produces two kinds of answer. The correctness answer says whether both the visible and hidden gates passed. The validity answer says whether the evidence is complete enough for primary analysis. A cell can therefore fail the task while remaining valid, or be invalid because the instrument could not establish a trustworthy result. The full suite is a regression signal beside those gates. It records whether the frozen suite passed and how many tests that passed on the pristine variant now fail. The hidden tier is the principal task-specific gate: its checks are selected by the task manifest, and a missing hidden scorer is recorded as an instrument failure rather than silently converted into a failed agent result.

The scoring ladder is visible tests, the full frozen suite, hidden levels, and an optional judge. The visible tests are the registered command that the run used for pass gating. The full suite is compared with a cached pass set from the pristine variant. Hidden levels are task checks that the agent does not receive. A judge is a language model used by a hidden check only when the task configuration permits it, and its scoring tokens are kept in a scoring ledger separate from the agent token ledger. The ladder is not a promise that every task has every level. The task manifest and the hidden scorer decide what can run.

Scoring is deliberately best effort for evidence capture, not permissive about missing correctness evidence. Test commands write their raw output before it is parsed. A command timeout is represented in its output and return data. A missing manifest, missing visible output, missing hidden output, missing hidden scorer, unavailable shell, unfinished run loop, missing token usage, or an unwaived runner halt can make the cell invalid. Evidence capture happens only after the summary is on disk and cannot alter that verdict. This position closes the execution loop and supplies the summary, raw files, timing roll-up, and evidence records read by aggregation, audit, the interface, and later review.

2. The reader’s map of the owning files

The principal owner is 5. Experiment/1. Harness/scripts/score_cell.py, a script of more than 50,000 bytes. Its own comments divide the relevant work into constants and helpers, visible and full-suite scoring, hidden scoring, invalidity evidence, and the final cell flow. Lines 1 to 180 contain imports, constants, path helpers, process execution and process-group termination. Lines 180 to 727 contain hidden-scorer configuration, baseline and handoff helpers, and the parsers used by the scoring steps. Lines 728 to 800 contain visible_step and full_suite_step. Lines 806 to 848 contain hidden_step and _write_scoring_ledger. Lines 852 to 1056 contain the shell evidence and invalidity contract, including collect_invalid_reasons. Lines 1059 to 1255 contain transcript model extraction, halt and waiver helpers, and reached-step support. Lines 1258 to 1400 contain score_cell, the main flow and summary writes. Lines 1403 to 1457 contain _refresh_stage_timings and best-effort evidence invocation. Lines 1460 to 1558 contain validity refresh. Lines 1597 to 1639 contain command-line parsing and main.

The function run_test_command in 5. Experiment/1. Harness/scripts/score_cell.py lines 99 to 160 is the common process helper. It runs a command with the scoring timeout, returns its status and text, and kills the process group when the timeout is reached. visible_step lines 728 to 759 writes the visible raw result and handles disagreement with the agent’s recorded outcome. full_suite_step lines 783 to 800 computes or loads the baseline pass set, runs the frozen suite, and counts introduced regressions. hidden_step lines 806 to 834 dispatches the task scorer and saves its raw result. _write_scoring_ledger lines 837 to 848 writes judge transcript rows as scoring external events. collect_invalid_reasons lines 976 to 1056 applies the validity contract. score_cell lines 1258 to 1400 owns orchestration, verdict fields, and both copies of the summary. _refresh_stage_timings lines 1403 to 1424 updates the timing roll-up. _capture_evidence_best_effort lines 1427 to 1457 invokes the separate capture program after scoring.

The evidence owner is 5. Experiment/1. Harness/scripts/capture_evidence.py, whose source is smaller than 50,000 bytes. Lines 1 to 74 contain imports, ports and configuration. Lines 75 to 178 contain Log, port and HTTP helpers, container proof, and demo-project seeding. Lines 180 to 285 contain capture_behaviour, including API and UI subprocesses, probes, and cleanup. Lines 286 to 388 contain capture assembly and manifest generation. Lines 390 to 441 contain capture and main. The companion 5. Experiment/1. Harness/scripts/evidence_shots.mjs is the screenshot command called by the UI branch; its structural role is the route-to-image boundary, while the Python script owns process setup and cleanup.

The post-experiment owner is 5. Experiment/1. Harness/scripts/rescore_tiered.py. Lines 1 to 57 contain path and condition helpers. Lines 58 to 140 contain cell discovery, one-cell hidden execution, and comparison with the recorded verdict. Lines 158 to 222 contain main, argument parsing, serial iteration, and the aggregate output. The audit owner is 5. Experiment/1. Harness/scripts/classify_instrument_errors.py. Lines 1 to 113 contain safe loading and parsing of final text and tier records. Lines 115 to 231 contain independent resolution and single-cell classification. Lines 233 to 271 contain audit aggregation and report formatting. Lines 273 to 293 contain main and its two output writes. The repair owner is 5. Experiment/1. Harness/scripts/repair_invocation_errors.py. Lines 1 to 42 contain scanning helpers, lines 43 to 85 contain loading, dumping and cell repair, and lines 87 to 139 contain run-set traversal and main.

The task-library contract is distributed between 5. Experiment/4. Task Library/task_manifest_schema.md, which defines the manifest fields that identify visible and hidden scoring, and 5. Experiment/4. Task Library/PROBE_AUTHORING.md, which governs authoring of hidden probes. The hidden-check convention is a task-local hidden_checks/checks.py selected from the task manifest. The judge switch is outside the task folder in 5. Experiment/1. Harness/config/hidden_judge.json, as the shared hidden-check documentation states. scoring_contract.json supplies the invalidity grounds consumed by the scorer. These are inputs to the main script rather than additional chapter-owned executables.

3. The inputs

3.1 Command-line arguments

ProgramArgument parser and lineArgumentDecision
score_cell.pymain lines 1597 to 1636--run-setSelects the run-set directory.
score_cell.pymain lines 1597 to 1636--cellSelects the cell to score.
score_cell.pymain lines 1597 to 1636--refresh-validityRecomputes validity without rerunning tests.
score_cell.pymain lines 1597 to 1636--dry-runSelects non-writing refresh behavior where supported by the dispatch.
score_cell.pymain lines 1597 to 1636--skip-evidencePrevents the post-verdict evidence subprocess.
score_cell.pymain lines 1597 to 1636--force-rescore-reachedAllows a reached chained step to be rescored instead of returning its existing summary.
rescore_tiered.pymain lines 158 to 222--run-setSupplies preserved cells as the input source.
rescore_tiered.pymain lines 158 to 222--taskSelects the task-local hidden checks.
rescore_tiered.pymain lines 158 to 222--stagedSupplies an alternative staged tree source.
rescore_tiered.pymain lines 158 to 222--onlyFilters discovered cells.
rescore_tiered.pymain lines 158 to 222--timeout-secondsBounds one hidden-check subprocess, defaulting to 900 seconds.
rescore_tiered.pymain lines 158 to 222--outSelects the aggregate rescore file.
capture_evidence.pymain lines 429 to 441--run-setSelects the run set containing the cell.
capture_evidence.pymain lines 429 to 441--cellSelects the cell whose evidence is captured.
capture_evidence.pymain lines 429 to 441--skip-uiOmits the UI server and screenshot branch.
capture_evidence.pymain lines 429 to 441--timeout-secondsBounds local server readiness and capture work.
classify_instrument_errors.pymain lines 278 to 293--run-sets-rootSelects the root scanned for instrument-error records.
classify_instrument_errors.pymain lines 278 to 293--json-outputSelects the audit JSON output.
classify_instrument_errors.pymain lines 278 to 293--markdown-outputSelects the human-readable audit report.
repair_invocation_errors.pymain lines 105 to 139--run-setSelects one run set to inspect.
repair_invocation_errors.pymain lines 105 to 139--all-run-setsSelects all available run sets.
repair_invocation_errors.pymain lines 105 to 139--applyChanges records rather than reporting proposed repairs.

The scorer’s ordinary caller is the cell runner, but the parser also supports validity refresh and manual reached-step handling. The rescorer requires either --run-set or --staged, and its parser refuses to proceed without a source. The evidence, audit and repair programs are separate command-line entry points, so their failures do not become a hidden test result unless the main scorer has already written that result.

3.2 Environment variables

VariableRead or set atDecision
FIVEB_HANDOFFscore_cell.py lines 383 to 391 and 630 to 650Carries the handoff path to a hidden scorer when the scorer uses it.
FIVEB_SKIP_EVIDENCEscore_cell.py lines 1438 to 1439Skips evidence capture when set.
FIVEB_CONDITIONscore_cell.py lines 506 to 508 and rescore_tiered.py lines 104 to 108Selects the cell condition for hidden checks.
FIVEB_WORKSPACEscore_cell.py lines 506 to 507Names the workspace presented to the hidden scorer.
FIVEB_BASELINE_TREEscore_cell.py lines 522 to 524 and rescore_tiered.py lines 104 to 108Names the pristine or baseline tree used by checks when one exists.
FIVEB_HIDDEN_TIER_OUTscore_cell.py lines 534 to 535 and rescore_tiered.py lines 104 to 108Names the hidden tier result destination.
PYTHONPATHinherited by rescore_tiered.py and read by the evidence subprocess environmentMakes the task and workspace imports resolve to the selected tree.
VITE_WS_URLcapture_evidence.py lines 351 to 356Gives the UI development server the local API WebSocket address.

The scorer sets the FIVEB_* values for a hidden subprocess rather than treating them as operator configuration. The rescorer builds a fresh environment for each preserved cell. The evidence program copies the process environment and adds the UI connection value. No network credential is an input to these local scoring programs.

3.3 Files read

File or directoryRead atDecision
cell_manifest.json or reproducibility_manifest.json_load_manifest, score_cell.py lines 1283 to 1285Supplies task directory, variant root, condition and registered commands.
workspace/score_cell.py lines 1311 to 1316 and hidden dispatch lines 650 to 750Is the agent’s tree under test.
scoring_outputs/score_cell.py lines 1269 to 1271 and step functionsHolds prior outputs and receives raw scoring files.
transcripts/pass*.stream.jsonlshell_evidence lines 919 to 970 and _transcript_modelsSupplies shell evidence, usage and model identity.
metrics.jsoncollect_invalid_reasons lines 1004 to 1016 and _refresh_stage_timings lines 1418 to 1424Establishes run completion and receives the timing roll-up.
operator_waiver.json_read_waiver and collect_invalid_reasonsCan excuse a specified runner halt.
task_manifest.json and hidden scoring specificationhidden configuration helpers and hidden_step lines 806 to 834Select the hidden scorer and its task inputs.
visible_tests.md or HOW_TO_VERIFY.mdvisible command selection in score_cell lines 1293 to 1305Supplies a registered visible command when the manifest does not.
baseline_pass_set.jsoncompute_or_load_baseline_pass_set lines 765 to 780Caches the pristine variant pass identifiers.
command_log.jsonl and container identifier filescapture_evidence.py lines 108 to 115Supply container proof.
evidence.jsoncapture_evidence.py lines 236 to 266Supplies API probes and UI routes when present.
hidden_checks/checks.pyrescore_tiered.py lines 38 to 44 and 95 to 123Is the task-specific preserved-tree scorer.
scoring_summary.json and hidden tier result snapshotsrescore_tiered.py lines 126 to 140 and classify_instrument_errors.py lines 31 to 58Supply comparison and audit evidence.

3.4 Network endpoints

The main scorer, the tiered rescorer, the instrument audit and the invocation repair program have no network endpoint. They run local commands or read local records. Evidence capture is different. _api_call in capture_evidence.py lines 165 to 191 calls the local API at 127.0.0.1:8791. The UI server is started on 127.0.0.1:5791, and evidence_shots.mjs is invoked against that local UI. _wait_http lines 154 to 163 polls these local addresses. These are temporary evidence services, not external experiment services.

4. The happy path in order

4.1 The reached-step decision

reached_step_summary in 5. Experiment/1. Harness/scripts/score_cell.py lines 1236 to 1255 begins the path by looking for a prior summary whose correctness gate is satisfied_by_start_tree. score_cell lines 1258 to 1268 returns that record unchanged unless --force-rescore-reached was supplied. A reached step is a chained task already satisfied by the starting tree, so rerunning a scorer that expects agent-created artefacts could replace a valid reached record with a misleading missing-output result. The flow-model nodes are scorer.reached_check and scorer.reached_kept. The ordinary batch driver does not call the scorer for this case, which leaves this branch for a manual rescore.

4.2 The prior-record preservation

_preserve_prior_scoring in 5. Experiment/1. Harness/scripts/score_cell.py lines 1199 to 1230 begins the ordinary path inside the preserve_prior_scoring timer at lines 1279 to 1282 and ends before the manifest is loaded. Existing scoring outputs are copied to dated prior files. The design reason in the surrounding comments is that the scorer must retain the earlier verdict before it overwrites the contract file, so a later audit can distinguish an original record from a rescored one. The flow-model node is scorer.preserve_prior_scoring.

4.3 The manifest and command selection

_load_manifest in 5. Experiment/1. Harness/scripts/score_cell.py lines 1280 to 1300 begins the input resolution, and the command selection in score_cell lines 1283 to 1307 ends it. The scorer prefers the manifest’s registered visible command, then its list of visible commands, then the task’s visible instructions, and finally an empty command list. It also obtains the variant root, task directory, condition, phase and full-suite command. The comments give the reason for using the registered command: the scorer must repeat the command used for pass gating, not every fenced command that the task author may have offered as an optional aid. A missing task directory is carried forward as a missing hidden scorer rather than treated as a hidden failure.

4.4 The visible suite

visible_step in 5. Experiment/1. Harness/scripts/score_cell.py lines 728 to 759 begins when the timer calls it at lines 1309 to 1311 and ends when it returns visible_pass, flaky_pass, the harness result, the agent result, and output paths. It runs the registered command, writes visible.final.txt before interpreting the result, and compares the fresh outcome with the agent’s recorded visible outcome. If they differ, it runs the bounded rerun set, writes each visible.rerun<N>.txt, and takes a majority. The comments give two reasons: raw output must survive parser failure, and a disagreement must be treated as possible flakiness rather than silently accepting either side. The flow-model node is scorer.score_visible_suite.

4.5 The full-suite regression gate

compute_or_load_baseline_pass_set in 5. Experiment/1. Harness/scripts/score_cell.py lines 765 to 780 begins the full-suite work, and full_suite_step lines 783 to 800 ends it. The pristine variant is run only when no baseline_pass_set.json exists. The workspace then runs the frozen full-suite command, whose raw text is written to full_suite.final.txt. The result includes the full-suite pass, the baseline path, and the count of identifiers that were in the pristine pass set but failed in the workspace. The comments say the baseline is computed once per variant and reused, which makes regressions comparable across cells and avoids giving each cell a different pristine measurement. When no variant root is available, score_cell lines 1312 to 1319 returns a not_run full-suite state.

4.6 The hidden tier

hidden_step in 5. Experiment/1. Harness/scripts/score_cell.py lines 806 to 834 begins when score_cell enters score_hidden_checks at lines 1320 to 1327 and ends with the hidden pass, missing-scorer flag, raw path and optional JSON path. The hidden dispatch uses the task’s configured scorer. The task library’s hidden_checks/checks.py convention keeps assertions outside the agent’s workspace view, while PROBE_AUTHORING.md governs how those probes are written and task_manifest_schema.md identifies the task scorer. The raw result is written to hidden.final.txt before its status is used. If the scorer is absent, hidden_pass is null and hidden_scorer_missing is true. The comment gives the design reason: a missing instrument must not look like a real agent failure. The flow-model node is scorer.score_hidden_checks.

4.7 The external scoring ledger

_write_scoring_ledger in 5. Experiment/1. Harness/scripts/score_cell.py lines 837 to 848 begins and ends the ledger step, called by score_cell lines 1329 to 1333 only when hidden scoring returned a transcript. It converts judge transcript rows into scoring_external events and writes scoring_ledger.jsonl. The comment gives the reason explicitly: judge tokens belong to the scoring operation and must never be appended to the agent’s token_ledger.jsonl. The flow-model node is scorer.score_write_ledger.

4.8 The validity and verdict calculation

collect_invalid_reasons in 5. Experiment/1. Harness/scripts/score_cell.py lines 976 to 1056 begins the validity decision, called at lines 1334 to 1336. It checks the manifest, visible output, hidden output, hidden scorer status, shell evidence, metrics completion, transcript usage and runner halt reasons, then reads any operator waiver. score_cell lines 1337 to 1371 combines the visible and hidden booleans for correctness, makes any nonempty invalidity list take precedence over pass or fail, and fills the five record states. A hidden scorer that is missing therefore gives the hidden and correctness states invalid, while an ordinary failed hidden check gives fail. The comments give the reason for each instrument guard: the summary must distinguish an unmeasured cell from an unsuccessful agent.

4.9 The summary and timing records

score_cell in 5. Experiment/1. Harness/scripts/score_cell.py lines 1347 to 1400 builds the summary and writes it first to the cell root and then to scoring_outputs/scoring_summary.json. The fields include per-pass gates, visible, hidden and full-suite outcomes, introduced regressions, validity, score states, invalid reasons, waiver information, transcript models, shell evidence and paths. _refresh_stage_timings lines 1403 to 1424 then reads the timing rows, replaces only stage_timings in metrics.json, and suppresses failures in that roll-up. The surrounding comment says scoring is often the slowest part of a cell and that a failed roll-up must not erase a complete line-by-line timing file. The flow-model nodes are scorer.verdict, scorer.summary_written, and scorer.stage_timings_refreshed.

4.10 The evidence capture

_capture_evidence_best_effort in 5. Experiment/1. Harness/scripts/score_cell.py lines 1427 to 1445 begins after the summary and timing work. _run_evidence_capture lines 1447 to 1457 starts capture_evidence.py with a 1,200 second subprocess bound and suppresses its failure into a stderr note. The command-line main function in score_cell.py lines 1597 to 1636 calls this helper after score_cell returns. The capture program’s capture function, in 5. Experiment/1. Harness/scripts/capture_evidence.py lines 390 to 428, collects container proof and behavioural evidence, then writes its manifest. The design reason in the scorer comment is ordering and independence: the verdict must already be durable, and evidence capture must not block or rewrite it. The flow-model terminal is scorer.finished, after the driver receives control.

4.11 The tiered rescore

task_dir in 5. Experiment/1. Harness/scripts/rescore_tiered.py lines 38 to 44 begins task lookup, discover lines 58 to 82 identifies preserved cells, and run_one lines 85 to 123 runs one task scorer. main lines 158 to 222 serially gathers the records and writes rescore_tiered.json. The program sets the condition and absolute workspace and baseline paths for each subprocess, reads its hidden tier result, adds cell metadata, and retains the original recorded verdict for comparison. The design reason is isolation from the agent run: changed hidden checks can be evaluated on preserved trees without spending agent tokens or changing the original summary.

5. The state machine

A verdict state is the value a scoring record can leave in one of its gate fields. pass means the relevant gate returned true. fail means it ran and returned false. not_run means the full suite had no variant root and was not attempted. invalid means the instrument could not support a primary-analysis conclusion. satisfied_by_start_tree means a chained step was already correct before an agent ran. The final overall field normally contains pass, fail, or invalid; the other values occur in component states or the reached-step record.

The following state diagram is the verdict state machine. It is hand-authored from score_cell.py lines 1262 to 1268, 1334 to 1371, and 1394 to 1397. The flow-model subgraph nodes represented by the decision and record transitions are scorer.reached_check, scorer.reached_kept, scorer.invalid_reasons, scorer.verdict, and scorer.summary_written.

stateDiagram-v2
    [*] --> not_run: component has no runnable variant
    [*] --> satisfied_by_start_tree: prior reached summary
    not_run --> invalid: missing run input or invalidity reason, score_cell 1317..1319 and 1334..1345
    not_run --> pass: component is advisory and correctness gates pass, score_cell 1337..1371
    not_run --> fail: component is advisory and correctness gate fails, score_cell 1337..1371
    satisfied_by_start_tree --> satisfied_by_start_tree: no force flag, score_cell 1262..1268
    satisfied_by_start_tree --> pass: forced scoring passes, score_cell 1337..1371
    satisfied_by_start_tree --> fail: forced scoring fails, score_cell 1337..1371
    satisfied_by_start_tree --> invalid: forced scoring finds an instrument reason, score_cell 1334..1345
    pass --> invalid: a validity reason is present, score_cell 1342..1345
    fail --> invalid: a validity reason is present, score_cell 1342..1345
    pass --> [*]: summary written, score_cell 1394..1400
    fail --> [*]: summary written, score_cell 1394..1400
    invalid --> [*]: summary written, score_cell 1394..1400

The scoring ladder is a separate hand-authored flowchart. It shows why a not_run full suite does not by itself make the cell invalid, while a missing hidden scorer does. Visible and hidden results form correctness; the full suite supplies regression information; the judge is nested inside hidden scoring and is not a second agent outcome.

flowchart LR
    A[Visible registered suite] --> B{Visible pass}
    B --> C[Full frozen suite]
    C --> D[Baseline pass set and regressions]
    D --> E[Hidden levels]
    E --> F{Deterministic hidden levels}
    F -->|unresolved and configured| G[Judge]
    F -->|resolved| H[Hidden pass or fail]
    G --> H
    H --> I{Validity contract}
    B --> I
    I -->|valid and both gates pass| J[overall pass]
    I -->|valid and a gate fails| K[overall fail]
    I -->|any invalid reason| L[overall invalid]

The full chapter 08 flow-model subgraph is the source of the next sequence diagram. Its node identifiers are scorer.reached_check, scorer.reached_kept, scorer.preserve_prior_scoring, scorer.score_visible_suite, scorer.score_full_suite, scorer.score_hidden_checks, scorer.score_write_ledger, scorer.invalid_reasons, scorer.verdict, scorer.summary_written, scorer.stage_timings_refreshed, and scorer.finished. The model records the main edges and their code ranges. The sequence participants below expand those nodes into programs, a workspace, test commands, hidden checks, records, and the local evidence services.

sequenceDiagram
    participant Runner as run_cell.py or operator
    participant Scorer as score_cell.py
    participant Manifest as cell_manifest.json
    participant Workspace as workspace/
    participant Visible as registered visible command
    participant Full as frozen suite
    participant Hidden as hidden_checks/checks.py
    participant Records as scoring outputs
    participant Evidence as capture_evidence.py
    participant API as 127.0.0.1:8791
    participant UI as 127.0.0.1:5791

    Runner->>Scorer: score_cell(run_set, cell_id), lines 1258..1400
    Scorer->>Records: preserve prior scoring, lines 1279..1282
    Scorer->>Manifest: load manifest, lines 1283..1307
    Scorer->>Visible: run registered command, visible_step 728..759
    Visible-->>Records: visible.final.txt and reruns
    Scorer->>Full: run frozen suite, full_suite_step 783..800
    Full-->>Records: full_suite.final.txt and baseline pass set
    Scorer->>Hidden: run task scorer, hidden_step 806..834
    Hidden-->>Records: hidden.final.txt and optional hidden JSON
    Scorer->>Records: write scoring ledger, lines 1329..1333
    Scorer->>Manifest: inspect transcript, metrics and waiver
    Scorer->>Records: write scoring summaries, lines 1394..1397
    Scorer->>Records: refresh metrics stage timings, lines 1403..1424
    Scorer->>Evidence: invoke after verdict, lines 1427..1457
    Evidence->>API: start, seed and probe locally
    Evidence->>UI: start and capture routes when enabled
    Evidence-->>Records: evidence files and manifest
    Scorer-->>Runner: summary and process completion

6. The sequence of one unit of work

One unit of work is one call that scores one cell. The sequence starts after the agent run has left its workspace and ends after the scorer returns its summary. The normal caller is the cell runner, while the same program can be called by a rescore operator. The following diagram is hand-authored from score_cell.py lines 1258 to 1457 and capture_evidence.py lines 390 to 428. The participants are the programs, local commands, files and temporary services that exchange the unit’s evidence.

sequenceDiagram
    participant Caller as cell runner or operator
    participant S as score_cell.py
    participant M as manifest files
    participant W as workspace
    participant V as visible tests
    participant F as full suite and baseline
    participant H as hidden scorer
    participant T as scoring records
    participant E as capture_evidence.py
    participant A as local API
    participant U as local UI

    Caller->>S: invoke one cell
    S->>T: copy prior scoring files
    S->>M: load manifest and commands
    S->>V: execute registered visible command
    V-->>T: visible.final.txt
    S->>F: execute frozen suite against W
    F-->>T: full_suite.final.txt and baseline_pass_set.json
    S->>H: execute hidden checks with workspace environment
    H-->>T: hidden.final.txt and hidden_tier_result.json
    S->>T: write scoring_ledger.jsonl when judge transcript exists
    S->>M: read metrics, transcripts and waiver
    S->>T: write cell and output summaries
    S->>T: rewrite metrics stage_timings
    S->>E: invoke evidence after verdict
    E->>A: start, seed and probe API
    E->>U: start and photograph UI when enabled
    E-->>T: evidence files and evidence_manifest.json
    S-->>Caller: summary and exit status

The ordering is material. Raw test output precedes parsing, the summary precedes evidence capture, and the timing roll-up follows the scoring stages. A failure in the last evidence branch therefore cannot erase or revise the verdict that the caller and aggregator already read.

7. The records

A scoring record is a file or table written by this mechanism and consumed later as evidence, a verdict, a timing input, or an audit input. The canonical summary belongs to this chapter. A reader that uses the summary for aggregation reads the cell-root copy, while tools that need raw outputs commonly read the copy in scoring_outputs.

RecordWriter and lineFields or contentsReaders
scoring_summary.json at the cell rootscore_cell.py, score_cell lines 1347 to 1397cell_id, passes, visible and hidden outcomes, hidden scorer flag, full-suite outcome, introduced regressions, flakiness, validity, score_states, invalid reasons, waiver, transcript models, shell evidence and pathsrun_batch.py and aggregate_metrics.py, plus the results interface where the cell summary is displayed
scoring_outputs/scoring_summary.jsonscore_cell.py, score_cell lines 1392 to 1397; repair_invocation_errors.py, repair_cell lines 64 to 85The same summary as the cell-root contract recordclassify_instrument_errors.py lines 148 to 231, rescore_tiered.py lines 126 to 140, repair scans, and audit readers
scoring_outputs/visible.final.txtscore_cell.py, visible_step lines 734 to 735Raw registered visible command outputAggregation and manual audit; the scorer checks that it exists at collect_invalid_reasons lines 983 to 984
scoring_outputs/visible.rerun<N>.txtscore_cell.py, visible_step lines 743 to 750Raw output for each flakiness rerunHuman audit and summary paths recorded by score_cell.py
scoring_outputs/full_suite.final.txtscore_cell.py, full_suite_step lines 790 to 793Raw full-suite outputAggregation and manual regression review
variant_root/baseline_pass_set.jsonscore_cell.py, compute_or_load_baseline_pass_set lines 765 to 780Variant name, command and sorted pristine pass identifiersLater full-suite scoring in full_suite_step, and downstream regression analysis
scoring_outputs/hidden.final.txtscore_cell.py, hidden_step lines 814 to 824Raw hidden scorer output, including an instrument error when the scorer is absentclassify_instrument_errors.py lines 71 to 79, the invalidity check at score_cell.py lines 985 to 990, and human audit
scoring_outputs/hidden.final.jsonscore_cell.py, hidden_step lines 825 to 833JSON form of hidden output when raw output parses as JSONHidden-tier analysis and instrument audit where present
scoring_outputs/hidden_tier_result.jsonhidden scorer subprocess through score_cell.py hidden dispatchTier outcome and, where applicable, resolved tier, first failure, judge information and raw detailclassify_instrument_errors.py lines 31 to 113, the results interface’s evidence views, and task analysis
scoring_outputs/hidden_tier_result.pass<N>.jsonhidden scorer subprocess during a pass-specific scoring runPer-pass hidden tier resultclassify_instrument_errors.py chooses the snapshot matching the final pass
scoring_ledger.jsonlscore_cell.py, _write_scoring_ledger lines 837 to 848One JSON line per judge scoring event, marked event_type scoring and category scoring_externalToken and cost analysis, not the agent token ledger
metrics.jsonscore_cell.py, _refresh_stage_timings lines 1403 to 1424; repair script lines 64 to 85Existing metric fields with refreshed stage_timingsaggregate_metrics.py, validity checks for run completion, and timing analysis
evidence/capture_log.txtcapture_evidence.py, Log and capture lines 75 to 83 and 390 to 428Timestamped capture notes and suppressed errorsAuditors and evidence debugging
evidence/container_evidence.jsoncapture_evidence.py, collect_container_proof lines 108 to 115Manifest container, billing, provider, pass container IDs, Docker invocations and notesAuditors and analysis scripts
evidence/api/<probe_name>.jsoncapture_evidence.py, _api_call lines 165 to 191 and capture_behaviour lines 236 to 365API request, response and parsed data or an errorAnalysts reviewing behavioural evidence
evidence/ui/*.pngevidence_shots.mjs, invoked by capture_evidence.py lines 347 to 369Screenshots of configured UI routesAnalysts and human reviewers
evidence/evidence_manifest.jsoncapture_evidence.py, capture lines 360 to 388Capture version, cell identity, behaviour summary and SHA-256 file hashesHarness verification and evidence review
rescore_tiered.jsonrescore_tiered.py, main lines 158 to 222Schema version, task, generation time and cell array containing preserved-tree outcomesRescore analysis and comparison with original summaries
audit JSON outputclassify_instrument_errors.py, main lines 278 to 288Source root, counts and records with evidence class, disposition, errors, resolution and conflictsData-integrity review and analysis preparation
audit Markdown outputclassify_instrument_errors.py, main lines 278 to 288Human-readable audit reportOperators and reviewers
repaired metrics.json and summariesrepair_invocation_errors.py, dump and repair_cell lines 60 to 85Corrected invocation-error fields when --apply is usedAggregation and later audit; the repair program itself reads its prior versions

The lineage is from execution evidence to the canonical summary, then to campaign tables. The raw files remain alongside the summary so that an aggregate boolean is not the only surviving proof. This diagram is hand-authored from the writer and reader relationships in the table. It is not a new flow-model claim.

flowchart LR
    W[workspace and transcripts] --> V[visible.final.txt]
    W --> F[full_suite.final.txt]
    B[pristine variant] --> P[baseline_pass_set.json]
    W --> H[hidden.final.txt]
    H --> R[hidden_tier_result.json]
    R --> L[scoring_ledger.jsonl when judged]
    V --> S[scoring_summary.json]
    F --> S
    R --> S
    M[metrics.json and waiver] --> S
    S --> C[aggregate_metrics and campaign tables]
    S --> I[classify_instrument_errors]
    H --> I
    R --> I
    S --> Q[results interface]
    W --> E[capture_evidence.py]
    E --> X[evidence_manifest and evidence files]
    X --> Q
    X --> C

The audit and repair records are deliberately downstream records. classify_instrument_errors.py never changes the historical verdict. It classifies conflicts, prior invalidity and independent deterministic resolution into an audit output. repair_invocation_errors.py changes records only when its apply flag is used, and its output remains identifiable as a repair rather than a fresh score.

8. The loops and the waits

The visible flakiness loop is in visible_step lines 743 to 752. Its bound is n_rerun, obtained by _flaky_rerun_n and defaulting to three. It runs only when the agent’s recorded visible result is not equal to the fresh harness result. It stops after the bounded reruns and chooses a strict majority, while flaky_pass records whether the rerun outcomes disagree. There is no retry of an ordinary failed visible suite when the agent and harness already agree.

run_test_command lines 99 to 160 waits for the child command until TEST_COMMAND_TIMEOUT_S, which the digest identifies as 900 seconds. At timeout it kills the process group and returns status 124 with the timeout marker in the captured text. The full suite and hidden scorer use this helper or their own subprocess bound, so an orphaned test process is not left behind by the normal timeout path.

parse_pytest_output and the transcript readers loop over finite lines. The visible and full-suite parsers stop at end of captured output. shell_evidence lines 921 to 970 loops over sorted transcript files and then their JSON lines and content blocks. Its stop condition is exhaustion of the final transcript set. It classifies a denial only when the session path marker is present and no shell execution is found; a permission-gate refusal is ignored as evidence that a shell ran.

refresh_run_sets, in score_cell.py lines 1550 to 1590, iterates through selected run sets and cells when validity is refreshed. It stops after the supplied paths and does not rerun tests. rescore_tiered.py main lines 158 to 222 iterates through the discovered cell list serially. The rescorer has no parallel retry because the digest records the performance constraint of the mounted filesystem. run_one lines 85 to 123 waits up to the --timeout-seconds value, default 900 seconds, and records a timed-out subprocess as not passed with a raw tail.

The evidence program has two HTTP readiness waits. _wait_http in capture_evidence.py lines 154 to 163 polls every three seconds until the API or UI responds successfully or the supplied timeout expires. The API and UI subprocesses are each given cleanup in capture_behaviour lines 236 to 365. The setup questionnaire in seed_demo_project lines 194 to 234 makes at most 30 API interactions, stopping when the server reports setup complete or when the loop ends. Screenshot subprocesses are given the capture timeout and then a 15 second termination wait before a kill, as implemented in the capture cleanup path.

The hidden tier itself may contain task-defined levels and a judge. Their internal bounds belong to the task-local hidden checker and the configured judge, not to this chapter’s scorer. The scorer’s bound is the enclosing hidden subprocess. Its stop condition is completion, timeout, or a returned instrument error. A result that cannot be decoded is retained as raw evidence rather than retried as a different verdict.

9. The guards and refusals

Guard or refusalLocationEffect
Reached summary without forcescore_cell.py lines 1262 to 1268Refuses to overwrite a satisfied_by_start_tree record and returns it.
Path traversal in a relative child_relative_child, score_cell.py lines 180 to 190Refuses a path that escapes the selected root before a hidden asset is staged.
Missing manifest_load_manifest and collect_invalid_reasons, lines 1283 to 1285 and 981 to 982Continues enough to write an invalid summary, with missing_reproducibility_manifest.
Missing visible outputcollect_invalid_reasons, lines 983 to 984Refuses primary validity by adding missing_visible_test_output.
Missing hidden outputcollect_invalid_reasons, lines 985 to 986Refuses primary validity by adding missing_hidden_scoring_output.
Hidden scorer missinghidden_step, lines 818 to 824, and invalidity lines 987 to 990Leaves hidden pass null and marks the cell invalid rather than failed.
Shell unavailableshell_evidence lines 887 to 970 and invalidity lines 1001 to 1002Refuses primary analysis when denial evidence exists without an executed shell call.
Run loop unfinishedcollect_invalid_reasons lines 1004 to 1016Adds run_loop_did_not_finish when metrics are absent.
Missing transcript usagecollect_invalid_reasons lines 1020 to 1037Adds missing_token_usage when the final result has no usage.
Runner halt_run_halt_reasons and collect_invalid_reasons lines 1049 to 1056Adds a run_halted:<status> reason unless the operator waiver covers it.
Operator waiver_read_waiver and collect_invalid_reasonsExcuses only the specified halt reason and records the waiver in the summary.
Baseline absencescore_cell lines 1312 to 1319Leaves the full-suite component not_run; it does not by itself refuse the cell.
Busy scoring port or absent workspacecapture_evidence.py lines 236 to 285Refuses behavioural capture but retains container proof and the existing verdict.
Node package lock mismatchcapture_evidence.py lines 300 to 356Refuses borrowing node_modules for the UI branch.
Missing tier resultclassify_instrument_errors.py lines 31 to 58Falls back to a root result or latest snapshot instead of assuming a result.
Conflicting summary fieldsclassify_instrument_errors.py lines 60 to 69 and 173 to 231Classifies the record as indeterminate and proposes quarantine.
Missing rescore sourcerescore_tiered.py parser lines 158 to 178Refuses to run without a run set or staged source.

These guards separate evidence refusal from task failure. The scorer can write a useful invalid record after a missing artifact, while capture and audit tools preserve what they can and identify what they could not establish.

10. Every unhappy path

Each row below is written in four parts. The first part states what the step is supposed to do. The second states why the implementation works that way. The third names the trigger and the record left. The fourth states the cost to the driver, scorer, aggregator and interface.

Trigger and code pathIntended step and design reasonRecord, status and exit codeCost to downstream readers
Missing manifest, _load_manifest lines 1280 to 1300Identify task, variant and commands. Continuing permits an auditable instrument result rather than an unrecorded crash.scoring_summary.json is written with missing_reproducibility_manifest; overall invalid; the scorer process normally returns success after writing.The driver receives an invalid score, aggregation excludes it from primary analysis, and the interface can show the reason.
Visible command timeout, run_test_command lines 99 to 160Run the registered visible suite to completion. The process group is killed so a timed-out test cannot continue changing the workspace.visible.final.txt contains the timeout marker and the command status is 124; the visible gate fails unless another invalid reason applies.The driver receives a completed scoring record, the scorer retains raw evidence, aggregation sees a failed gate, and the interface can link the output.
Full-suite timeout or command failure, full_suite_step lines 783 to 800Compare the workspace with the frozen regression suite. Raw output is saved before parsing to preserve the diagnostic.full_suite.final.txt is written; full_suite_pass is false and regressions are counted from parsed identifiers where possible.The driver still receives a summary, the scorer can compute correctness from visible and hidden gates, the aggregator receives an advisory regression failure, and the interface displays the component state.
No task directory or no executable hidden scorer, score_cell lines 1320 to 1327 and hidden_step lines 806 to 834Evaluate the hidden gate. A missing evaluator must be distinguishable from a failing agent.hidden.final.txt is written when the step runs, hidden_pass is null, hidden_scorer_missing is true, and scoring_summary.json is invalid.The driver gets an invalid result rather than a false fail, aggregation excludes the cell, and the interface can identify an instrument error.
Hidden scorer returns malformed or instrument-error output, hidden dispatch and hidden_step lines 814 to 833Preserve the scorer’s own explanation while preventing it from becoming a behavioural failure.Raw hidden output and any valid JSON copy remain; the summary records the missing or unusable scorer state and an invalid reason.The scorer cannot claim task correctness, the aggregator quarantines the cell, and the interface has raw material for audit.
Shell unavailable, shell_evidence lines 887 to 970Determine whether the agent had the command-line capability needed to inspect and test its workspace. The rule uses transcript evidence rather than date inference.scoring_summary.json contains shell evidence and shell_unavailable; overall is invalid.The driver finishes normally, the scorer preserves the counts, the aggregator excludes the cell, and the interface can explain the exclusion.
Run loop did not finish, collect_invalid_reasons lines 1004 to 1016Establish that the agent’s attempt loop reached its own completion point. Metrics absence means the recorded attempts may be partial.No new metrics are invented; the summary contains run_loop_did_not_finish and is invalid.The driver sees an invalid score, the aggregator excludes partial work, and the interface can distinguish it from an agent failure.
Missing token usage, collect_invalid_reasons lines 1020 to 1037Preserve the token accounting needed to interpret a cell. The scorer refuses to infer usage.The existing transcript remains, and the summary contains missing_token_usage; status is invalid with the normal scoring process exit.Aggregation excludes an unaccountable cell, while the interface and audit retain the transcript path.
Runner halt, _run_halt_reasons and waiver handling lines 1049 to 1056Carry a run-cell protocol halt into scoring. The waiver is explicit so an operator decision cannot be mistaken for an ordinary pass.Without a matching waiver, run_halted:<status> is in invalid_reasons; with one, the reason is excused and the waiver is recorded.The driver receives either invalid or counted data, the aggregator follows valid_for_primary_analysis, and the interface can show the waiver.
Evidence port already bound or workspace absent, capture_evidence.py lines 236 to 285Capture behavioural evidence without disturbing scoring. Skipping an unavailable branch is safer than taking over another service.capture_log.txt records the note; container proof may still be written; the verdict and exit status from scoring are unchanged.The driver and scorer keep their result, aggregation sees no changed score, and the interface has incomplete evidence rather than a false behavioural record.
API or UI startup failure, capture_behaviour lines 236 to 365Start temporary local services and collect probes and images. Suppression keeps a diagnostic tool from becoming an experiment gate.Server logs and capture_log.txt receive the failure; the evidence manifest may be partial or absent; capture returns without changing scoring.No effect on the driver verdict or aggregation; the interface can show missing evidence and reviewers can inspect logs.
API probe or screenshot failure, _api_call lines 165 to 191 and evidence_shots.mjs invocationPreserve each available observation while continuing to later probes. Independent probe files make one route failure local.The probe JSON contains an error or the screenshot is absent; later records can still be written.The scorer and aggregator are unaffected, while the interface and auditor see a partial evidence set.
Rescore timeout or nonzero hidden subprocess, rescore_tiered.py lines 85 to 123Re-evaluate a preserved workspace without hiding the fact that the new check did not pass.hidden_tier_result.json records passed false, an absent or captured exit code, seconds and raw tail; aggregate rescore records the cell.The original summary is not changed, so the primary aggregator and interface remain stable; rescore readers see a failed or timed-out new evaluation.
No rescore cells or missing task checks, task_dir, discover and main lines 38 to 82 and 158 to 222Establish a meaningful rescore population. Refusing avoids writing an apparently complete empty analysis.Missing checks raises the task lookup error; no-cell discovery returns one; no per-cell record is promised.The operator receives a non-success command result, while existing scoring, aggregation and interface data remain untouched.
Invalid audit JSON or missing tier snapshot, classify_instrument_errors.py lines 17 to 58Classify instrument errors using the strongest available record. Fallbacks prevent one corrupt auxiliary file from hiding a historical error.load_json returns none and the classifier falls back; audit JSON and Markdown still contain the classification when enough evidence exists.No verdict changes, aggregation is not rewritten, and the interface is unaffected; reviewers receive an indeterminate or fallback audit record.
Conflicting cell and output summaries, summary_conflicts and classify_record lines 60 to 231Detect data-integrity disagreement before proposing disposition. A conflict cannot safely be resolved by choosing one copy.Audit output marks the record indeterminate and proposes quarantine. The historical summary remains unchanged.The aggregator must exclude or quarantine the record, while the interface can continue to display the original and the audit explains the conflict.
Repair scan finds a repairable invocation error, repair_cell lines 64 to 85Restore the historical record shape without silently changing records during inspection. The apply flag is required for mutation.Without --apply, no record is changed. With it, repaired summaries and metrics are dumped and the repair report is printed.The driver is not rerun, the scorer is not rerun, aggregation reads the repaired record on later runs, and the interface sees the new file after the operator action.

The scorer’s command-line main function in 5. Experiment/1. Harness/scripts/score_cell.py lines 1597 to 1636 returns exit code 0 for a valid summary and exit code 2 when score_states.overall is invalid. Thus the scorer rows above that end in invalid have process status 2, while valid pass and fail rows have status 0. A test timeout is stored as child status 124 but does not itself determine the scorer process status. Evidence capture failures are suppressed by the evidence subprocess and do not change the already returned scoring status. The tiered rescorer returns 1 for no matching cells and records per-cell timeout or nonzero child status in hidden_tier_result.json; its ordinary aggregate completion is 0. The instrument audit writes its two outputs and returns 0, while repair mutations are controlled by --apply and are not a new scoring exit path.

11. The metrics produced or fed

The primary correctness fields are visible_pass and hidden_pass in scoring_summary.json. They are produced by visible_step and hidden_step, then combined by score_cell into score_states.correctness_gate and score_states.overall. full_suite_pass and regressions_introduced come from full_suite_step. They are advisory to the correctness gate because the full suite can be not_run, while the visible and hidden gates determine correctness when the cell is valid.

valid_for_primary_analysis is the inverse of the invalidity list and is produced by score_cell after collect_invalid_reasons. invalid_reasons, operator_waiver, shell_evidence, and models_in_transcript carry audit detail rather than a task-quality score. passes is the per-pass gate array consumed by aggregation for success-pass calculations. The path fields connect scalar columns back to raw files.

flaky_pass is produced by the visible rerun loop. It is advisory because it describes disagreement across repeated observations rather than replacing the visible result. baseline_pass_set becomes the source for the count regressions_introduced. The scorer writes stage_timings through _refresh_stage_timings; timing aggregation reads it to expose scoring cost beside agent execution cost. The scoring ledger supplies external judge token rows to cost analysis and is intentionally not folded into the agent token ledger.

Evidence fields are not correctness columns. Container identity, API responses, UI images, hashes and capture behaviour support later audit and interpretation. A failed evidence branch therefore produces missing or partial evidence rather than invalid in the scoring summary. The audit program’s evidence class, proposed disposition, independent resolution and conflict fields are analysis annotations derived from the scoring outputs, not replacements for the original verdict.

12. The tests

The source inventory reviewed for this chapter did not make a mapping from each unhappy path in section 10 to a named test file. The mechanism has code-level seams and the task library has hidden checks, but this review did not verify a complete test-to-exit-path matrix. Accordingly, no claim is made here that every timeout, missing artefact, waiver, capture failure, rescore failure and repair branch has a dedicated automated test. The absence of that mapping is itself the limit of this section.

The visible and hidden task checks exercise product behaviour, not every harness refusal. score_cell.py contains the most direct test seams through run_test_command, visible_step, full_suite_step, hidden_step, collect_invalid_reasons, and _refresh_stage_timings. capture_evidence.py similarly exposes local seams for port checks, HTTP readiness, API calls and subprocess cleanup. rescore_tiered.py, classify_instrument_errors.py, and repair_invocation_errors.py have deterministic file-oriented entry points. Their existence does not establish coverage of every path listed above.

13. The dated incidents that shaped the code

The comments in score_cell.py carry several dated rulings that explain guards which might otherwise look arbitrary. The shell-unavailable rule at lines 852 to 876 records the 2026-09-03 decision that a container whose command tool could not write its session directory did not measure the intended agent condition. The shell evidence comment at lines 892 to 914 records why a permission-gate refusal is not counted as a shell denial and why a cell that issued no shell call is not judged shell-less.

The same 2026-09-03 ruling appears in collect_invalid_reasons lines 992 to 1001 and lines 1039 to 1047. The former makes the shell defect invalid for primary analysis. The latter records the decision not to exclude a cell merely because more than one model appears in its transcript. The models remain visible in models_in_transcript so the record is auditable without treating an agent’s helper model as an instrument error.

The run-loop guard at score_cell.py lines 1004 to 1016 carries the example of run set 017, where a large failing-test message made a later retry prompt exceed the command argument limit. The guard prevents the partial record from entering analysis as an ordinary agent failure. The runner-halt guard at lines 1049 to 1056 carries the 2026-08-15 ruling after a protocol halt was once counted as valid, and the waiver mechanism is the operator-controlled exception to that rule.

The registered-command rule in score_cell.py lines 1293 to 1299 carries the 2026-07-26 operator deviation. It prevents optional fenced commands, including a pre-existing end-to-end failure not registered by the task, from changing the measured gate. The reached-step preservation at lines 1262 to 1268 carries tracker 2.35 and protects a chained step whose starting tree already satisfies the task. These comments are design history attached to active guards, not evidence that the guards are accidental.

14. The weakest claim, what was not checked, and the token line

The weakest claim is that the listed downstream readers and campaign columns are complete for every historical run. The code and digest verify the scorer’s canonical writes and several named readers, but this review did not inspect every aggregator query, every task-local hidden-check schema, or the complete automated test inventory. The state and sequence diagrams identify the chapter 08 model nodes and use hand-authored expansions where the model does not specify file-level detail. The token line is:

Model: openai/gpt-5.6-luna via OpenRouter; tokens: see the job ledger