Chapter 2.9 follows chapter 2.8, The model-serving routes, and precedes chapter 2.10, The harness test suite, within part 2, the harness.
The configuration set governs run limits, token budgets, costs, scoring, reproducibility, model routing, and agent handoff. The files reside in the harness configuration directory.
2.9.1 Configuration file roles of the measurement harness
The file run_defaults.json governs default limits, the preset ceilings on attempts; token budgets, the allowed token amounts; and the cost model, the rules for assigning costs. The file scoring_contract.json governs score states, the labels that describe outcomes; invalidity conditions, the reasons a result cannot be used; full-suite regression, a comparison of the post-fix workspace with the variant baseline to show whether the fix introduced regressions; and flaky reruns, repeated test runs used to resolve disagreement between the visible-test outcome produced by the agent and the visible-test outcome produced by the harness. The file hidden_judge.json governs the language-model judge used as a last resort. The file reproducibility_schema.json governs the required fields of every manifest. The file model_matrix.json governs model-to-route selection, the choice of model for each route. The file agent_handoff.schema.json governs the handoff record, the record passed between stages.
source: 5. Experiment/1. Harness/config/run_defaults.json lines 3 to 14
Configuration files and their governing responsibilities.
2.9.2 Run limits and budgets
The pass limit is five attempts for a cell.
source: 5. Experiment/1. Harness/config/run_defaults.json line 3 (key max_passes)
The paid route has a token budget of two and a half million tokens across passes, counted to success.
source: 5. Experiment/1. Harness/config/run_defaults.json line 4 (key token_budget_per_cell)
source: folder 5. Experiment/0. Plan section (d) line 112
The measured local route has a token budget of thirty million tokens because its provider reports no cache fields.
source: 5. Experiment/1. Harness/config/run_defaults.json line 6 (key token_budget_per_cell_dgx_claude)
The flaky re-run count is three when the agent-run visible-test result disagrees with the harness-run result.
source: 5. Experiment/1. Harness/config/run_defaults.json line 9 (key flaky_rerun_n)
2.9.3 Score states and invalidity conditions
The score states are pass, fail, not_run, and invalid.
source: 5. Experiment/1. Harness/config/scoring_contract.json keys score_states and invalid_if
The invalidity conditions are missing_visible_test_output, missing_hidden_scoring_output, missing_token_usage, missing_reproducibility_manifest, multi_model_transcript, hidden_scorer_missing, and shell_unavailable.
source: 5. Experiment/1. Harness/config/scoring_contract.json keys score_states and invalid_if
A full-suite regression records whether the comparison introduced regressions. A flaky rerun resolves disagreement between the agent-run and harness-run visible-test outcomes by using the majority outcome.
source: 5. Experiment/1. Harness/config/scoring_contract.json keys full_suite_regression and flakiness_rerun
2.9.4 Cost model of the measurement harness
The cost model assigns 0.70 dollars to a successful cell, 2.85 dollars to a failed cell, and 5.00 dollars to the worst case.
source: 5. Experiment/1. Harness/config/run_defaults.json lines 10 to 14
The model matrix supplies the route and model mapping used with these defaults. The harness uses that mapping when it calculates cell cost.
2.9.5 Language-model judge of the measurement harness
The language-model judge is a last-resort evaluator for an outcome that automated tests leave inconclusive. Its schema is 5b.hidden_judge.v001.
source: hidden_judge.json schema 5b.hidden_judge.v001
The judge evaluates the cell outcome from the hidden probes. The hidden judge configuration keeps this evaluation separate from the automated scoring contract.
2.9.6 Manifest fields of the measurement harness
A manifest, a record of a cell’s metadata used to reproduce its run, is checked against reproducibility_schema.json. The schema requires the fields needed to identify the cell and preserve its reproducibility record. The harness checks every manifest before filing the cell.
source: reproducibility_schema.json
2.9.7 Migrating text of the measurement harness
Page 07 line 163 cites the defaults file from this entry.
source: 0. The Website/R4 Measurement of one cell.md line 163