Chapter 2.1 opens part 2, the harness, after part 1, the reading path, and precedes chapter 2.2, The cell pipeline from staging to scoring.

2.1.1 Harness layers for the two program layers of the measurement harness

The harness, a set of programs that prepares and measures a cell, is organized into two functional layers with separate roles: the build layer, the part that prepares the project corpus and its artifact variants, and the measurement layer, the part that places an agent in a variant, applies a seeded task, records the run, and measures the result. The layers keep their token ledgers separate.

2.1.2 Folder map for the directory layout of the measurement harness

The harness folder contains 153 files, of which four are cache files. source: operations/site-ia/A1-reverse-index.json

The folder map, the directory layout showing what the folder contains, lists the project readme, the implementation specification, configuration, container material, scripts, and prompts. The scripts directory contains the programs used by a measured run. Configuration and container files control the environment in which those programs operate.

Harness folder map.

2.1.3 Cell programs of the measurement harness

The cell pipeline begins with run_cell.py. It resolves the task and variant, stages the working copy, applies the seeded defect, runs the agent, and records the transcript. It leaves the transcript, patch set, and cell manifest in the cell folder.

score_cell.py follows the agent run. It runs the visible tests, the regression suite, and the hidden scorer against the finished tree. It leaves the scoring summary and raw test outputs in the cell folder.

aggregate_metrics.py follows scoring. It reads the run ledgers and scoring summaries, then writes the aggregate metric record for the cell and campaign. The batch runner calls these programs in this order and preserves the cell outputs for later review.

2.1.4 Isolation boundaries of the measurement harness

The isolation contract, the rules that keep the agent and scorer apart, governs the measured cell. It describes the behaviour required between the agent under test and the host-side scoring machinery. The agent receives the task brief and visible test instructions. Hidden checks and their raw output remain outside the agent workspace.

Isolation boundaries between the agent and the scorer.

The isolation contract has eight sections. source: the contract’s heading lines above

The blind boundary, the separation that keeps hidden material away from the agent, keeps hidden check source, hidden check directories, hidden output, and other cell state outside the agent’s reach. Host-only scoring runs outside the agent container against the finished tree. The hidden judge receives behavioural evidence and cannot inspect source code or tests. The handoff model carries the agent’s declared candidates without hidden material.

A scorer reads only the build handed to the agent, the frozen starting point, and earlier steps in the same chain. It does not read a sibling workspace or another condition’s scoring output. Visible tests remain available to the agent. The regression suite and evidence capture provide later checks and records without changing the verdict.

2.1.5 Command contract of the measurement harness

The command contract, the agreed interface for invoking harness programs, covers building, running, scoring, aggregation, monitoring, reporting, analysis, publishing, preflight checks, and repair operations. Each command receives the inputs specified by the harness and writes the outputs expected by the next stage.

2.1.6 Container contract of the measurement harness

The container contract governs the image, available shell, execution context, project manifest, agent roles, ledgers, trace records, handoff material, and publication locks. It keeps the agent workspace distinct from host-side scoring and preserves the records needed to reconstruct a cell.

2.1.7 Guard mechanisms, controls that protect the harness of the measurement harness

Guard mechanisms, the controls that protect the harness from invalid access and disclosure, protect staging, admission, parity, visibility boundaries, evidence, handoff data, and the cell record. They reject disclosures in task material, keep hidden checks out of the workspace, validate handoff shape, and preserve evidence from the run. A stopped or failed batch still receives its closeout records.

2.1.8 Cell sequence of the measurement harness

The cell scroll sequence shows the agent entering the staged workspace, making the change, running visible checks, handing off its result, and passing the finished tree to scoring. The added stills and the caveat correction are recorded in the asset review.

Cell scroll sequence from staging to recorded results.

Previous: Measurement of one cell | Next: The cell pipeline from staging to scoring
source: operations/site-ia/A6-page-briefs.md, section ### H0.