Chapter 1.2 follows chapter 1.1, The documentation experiment and its present status, and precedes chapter 1.3, The measuring instrument, within part 1, the reading path.

1.2.1 Conditions under which the documentation claim can fail

The experiment tests a narrow claim about whether targeted documentation helps Claude Code, the coding agent used for the maintenance tasks, complete work in a codebase. The claim applies to a pocket, a set of task conditions in which a rule spans modules, the prompt leaves that rule unstated, search is costly, and the recorded rule can prevent a repair attempt or a failed check. The earlier campaign, the preceding campaign that tested a broader documentation claim, rejected that broad claim and supplied the pocket conditions for this test. source: operations/site-ia/A6-page-briefs.md line 93

The claim is falsifiable. It fails when the documented arm costs more tokens than the control under the pocket conditions. It also fails when the documented arm does not preserve the expected task outcome under those conditions. A token, one unit of model-use accounting, is the measured cost under the accounting rule. source: operations/site-ia/A6-page-briefs.md line 93

1.2.2 Earlier campaign evidence of the documentation experiment

The earlier campaign measured 117 comparable paired rows of token cost. A paired row compares the same maintenance task on two versions of the codebase. Only 26 rows showed a token saving for the documented arm.

source: 5. Experiment/9. Archive/campaign-5a/Research Initial Findings.md line 36

The aggregate result rejected a general win while leaving the pocket as a testable narrower claim. An audit record on page 01 also carries a seventeen-unit count without a named file in that record. The task index supplies the source for that count.

source: 4. Task Library/task_index.csv

1.2.3 Comparison arms for the documentation treatment

A variant, a frozen build tree used in the experiment, is one version of the same software. An arm, meaning one experimental version in a comparison, is described with the account. The three arms descend from one locked specification. Their differences concern how the code was produced and which documentation accompanies it. The variant trees are recorded in the variant library, with their structural differences described in Differences between the variant trees. source: operations/site-ia/A6-page-briefs.md line 93

The documented arm carries the full documentation layer. The stripped arm is derived from it by removing comments, docstrings, and documentation-only files while retaining executable code. The control is an independent specification-only build whose construction is documented in the control construction record. source: operations/site-ia/A6-page-briefs.md line 93

A dose, the amount of documentation carried by a build, describes the profiles derived from the documented build. The stripped profile has no documentation layer, while the documented profile carries the full layer. The profiles therefore come from one build lineage, with the control kept as an independent arm.

1.2.4 Comparison estimands of the documentation experiment

The experiment has three comparisons, each a named contrast between arms. The documented arm against the control is the confirmatory comparison and estimates the effect of the full intervention against unstructured practice. The documented arm against the stripped arm is a registered secondary comparison and isolates the documentation layer over shared executable code. The stripped arm against the control is a registered secondary comparison and describes the structure component.

source: 6. Metrics/statistical_analysis_plan.md line 59

The arm codes are the short labels used in the experiment account. A cell, one task execution under one experimental condition, belongs to a run set, a group of cells collected under a shared comparison and schedule, recorded in the run structure. source: operations/site-ia/A6-page-briefs.md line 93

1.2.5 Experimental controls of the documentation experiment

A seed, a random starting state for an agent run, is assigned according to the registered schedule. Claude Code runs under the same task, environment, budget, and measurement rules across the relevant arms.

A gate, a pre-run check that permits a cell to start only when its requirements hold, protects the run set from an unready environment or an unavailable budget. A freeze, a locked state of the codebase and environment, keeps the measured starting state fixed. Parity repair, a repair that restores an arm to specification parity, is recorded as a change to the control history and is not treated as documentation evidence.

The fairness rules hold task identity, software specification, environment, measurement procedure, and contamination controls constant. They keep the three comparisons interpretable. Exclusions cover failed gates, budget failures, contamination, and loss of comparability. Exclusion is a finding because it records where the design could not support an estimate.

1.2.6 Falsification record of the documentation experiment

The falsification record joins the claim, the pocket conditions, the arm construction, and the registered comparisons. The confirmatory contrast is documented against control. The two remaining contrasts are secondary. Results that fail a gate or cease to be comparable remain exclusions rather than silent repairs to the evidence set.

A decision path showing the pocket conditions, the documented arm, the control, and the outcomes that would falsify the claim Failure conditions for the claim.

A decomposition of the documented against control, documented against stripped, and stripped against control comparisons The three comparison decomposition.

A view of the shared task, environment, budget, seed, and contamination rules across the arms Fairness rules across the arms.

A verdict diagram showing the earlier campaign result and the narrower claim that remains source: operations/site-ia/A6-page-briefs.md line 97 The earlier campaign verdict.

A diagram showing the conditions that narrow the broad claim to the pocket Conditions for the narrow claim.

A diagram showing the design choices that distinguish the later experiment from campaign 5a source: operations/site-ia/A6-page-briefs.md line 97 Design choices after campaign 5a. source: operations/site-ia/A6-page-briefs.md line 97

A diagram showing the relationship between an old idea and a new reader of the codebase The old idea and the new reader.

A diagram contrasting pointer-based orientation with repository search Pointer-based orientation and search.

A diagram showing the conditions that make the test believable Conditions for a believable test.

Previous: The documentation experiment and its present status | Next: The measuring instrument
source: operations/site-ia/A6-page-briefs.md line 93