Chapter 1.2 follows chapter 1.1, The documentation experiment and its present status, and precedes chapter 1.3, The measuring instrument, within part 1, the reading path.
1.2.1 Conditions under which the documentation claim can fail
The experiment tests a narrow claim about whether targeted documentation helps Claude Code, the coding agent used for the maintenance tasks, complete work in a codebase. The claim applies to a pocket, a set of task conditions in which a rule spans modules, the prompt leaves that rule unstated, search is costly, and the recorded rule can prevent a repair attempt or a failed check. The earlier campaign, the preceding campaign that tested a broader documentation claim, rejected that broad claim and supplied the pocket conditions for this test. source: operations/site-ia/A6-page-briefs.md line 93
The claim is falsifiable. It fails when the documented arm costs more tokens than the control under the pocket conditions. It also fails when the documented arm does not preserve the expected task outcome under those conditions. A token, one unit of model-use accounting, is the measured cost under the accounting rule. source: operations/site-ia/A6-page-briefs.md line 93
1.2.2 Earlier campaign evidence of the documentation experiment
The earlier campaign measured 117 comparable paired rows of token cost. A paired row compares the same maintenance task on two versions of the codebase. Only 26 rows showed a token saving for the documented arm.
source: 5. Experiment/9. Archive/campaign-5a/Research Initial Findings.md line 36
The aggregate result rejected a general win while leaving the pocket as a testable narrower claim. An audit record on page 01 also carries a seventeen-unit count without a named file in that record. The task index supplies the source for that count.
source: 4. Task Library/task_index.csv
1.2.3 Comparison arms for the documentation treatment
A variant, a frozen build tree used in the experiment, is one version of the same software. An arm, meaning one experimental version in a comparison, is described with the account. The three arms descend from one locked specification. Their differences concern how the code was produced and which documentation accompanies it. The variant trees are recorded in the variant library, with their structural differences described in Differences between the variant trees. source: operations/site-ia/A6-page-briefs.md line 93
The documented arm carries the full documentation layer. The stripped arm is derived from it by removing comments, docstrings, and documentation-only files while retaining executable code. The control is an independent specification-only build whose construction is documented in the control construction record. source: operations/site-ia/A6-page-briefs.md line 93
A dose, the amount of documentation carried by a build, describes the profiles derived from the documented build. The stripped profile has no documentation layer, while the documented profile carries the full layer. The profiles therefore come from one build lineage, with the control kept as an independent arm.
1.2.4 Comparison estimands of the documentation experiment
The experiment has three comparisons, each a named contrast between arms. The documented arm against the control is the confirmatory comparison and estimates the effect of the full intervention against unstructured practice. The documented arm against the stripped arm is a registered secondary comparison and isolates the documentation layer over shared executable code. The stripped arm against the control is a registered secondary comparison and describes the structure component.
source: 6. Metrics/statistical_analysis_plan.md line 59
The arm codes are the short labels used in the experiment account. A cell, one task execution under one experimental condition, belongs to a run set, a group of cells collected under a shared comparison and schedule, recorded in the run structure. source: operations/site-ia/A6-page-briefs.md line 93
1.2.5 Experimental controls of the documentation experiment
A seed, a random starting state for an agent run, is assigned according to the registered schedule. Claude Code runs under the same task, environment, budget, and measurement rules across the relevant arms.
A gate, a pre-run check that permits a cell to start only when its requirements hold, protects the run set from an unready environment or an unavailable budget. A freeze, a locked state of the codebase and environment, keeps the measured starting state fixed. Parity repair, a repair that restores an arm to specification parity, is recorded as a change to the control history and is not treated as documentation evidence.
The fairness rules hold task identity, software specification, environment, measurement procedure, and contamination controls constant. They keep the three comparisons interpretable. Exclusions cover failed gates, budget failures, contamination, and loss of comparability. Exclusion is a finding because it records where the design could not support an estimate.
1.2.6 Falsification record of the documentation experiment
The falsification record joins the claim, the pocket conditions, the arm construction, and the registered comparisons. The confirmatory contrast is documented against control. The two remaining contrasts are secondary. Results that fail a gate or cease to be comparable remain exclusions rather than silent repairs to the evidence set.
Failure conditions for the claim.
The three comparison decomposition.
Fairness rules across the arms.
source: operations/site-ia/A6-page-briefs.md line 97
The earlier campaign verdict.
Conditions for the narrow claim.
source: operations/site-ia/A6-page-briefs.md line 97
Design choices after campaign 5a.
source: operations/site-ia/A6-page-briefs.md line 97
The old idea and the new reader.
Pointer-based orientation and search.
Conditions for a believable test.