Chapter 1.8 follows chapter 1.7, Findings of the shakedown batches, and precedes chapter 1.9, Pending freezes and registrations, within part 1, the reading path.

The experiment has limits even when its procedures run as specified. These limits concern the strength of the statistical conclusion, the causal interpretation, the meaning of the measures, and the reach of the findings. The page also records the limits created by a single model on one product, run-to-run scatter, and conflicting records.

The Operator Decision Sheet, the governing record for experimental choices and unresolved risks, holds the decision records for these confounds. A confound is a hidden variable that could account for an observed result. The run-set timing confound is the risk that timing or ordering differs with treatment. The control-build model confound is the risk that the control uses a different model or model context. A matched pair is two runs sharing a task and seed so their treatment contrast is direct. The letter k is the repetition count behind a comparison. A container is a sealed execution environment isolated from the host. Parity repair is restoration of the control build to the documented specification.

The researcher’s six roles, task selector, prompt designer, implementer, operator, analyst, and decision owner, place selection, execution, interpretation, and governance in the same person.

source: 5. Experiment/0. Plan/Appendix E, Protocol.md, section 4

1.8.1 Validity threat families of the documentation experiment

A validity threat is a reason a result might not mean what it appears to mean. Conclusion validity concerns noise and statistical support. Internal validity concerns whether the treatment caused the result. Construct validity concerns whether the measures represent the intended properties. External validity concerns how far the result travels beyond the tested setting. Each threat has a countermeasure and a residual.

A family map filing threats under conclusion, internal, construct, and external validity, with a countermeasure and residual for each family.

Validity families and their remaining limits.

A flow diagram showing a validity threat, its countermeasure, and its residual, with the residual connected to an owning decision item.

Threat, countermeasure, and residual.

The conclusion-validity limit is the run-to-run scatter visible in paired responses. Pairing, aggregation, and correction reduce misreading of noise, while the observed scatter still limits the claims supported by the comparisons.

The internal-validity limits include the control-build model confound, the run-set timing confound, and the scope of the fingerprint guard. Staging gates, run guards, sealed containers, and parity repair address parts of these risks. The remaining control mismatch is held for the redesign described in the task-selection record.

The construct-validity limit includes the difference between curated trace documentation and ordinary comments in the control. Hidden probes, transcript-based token accounting, and restricted feedback address some measurement risks. Difficulty rating and prompt sensitivity remain limited.

The external-validity limits include task selection by a researcher aware of the treatment, a single command-line agent, model drift under a pinned name, and the narrow anchor of one model on one product.

A bias map showing how task knowledge can influence selection and how task ordering, review, and registration constrain that influence.

Selection bias and its countermeasures.

A role map showing the researcher occupying the experiment's roles and the resulting risks for selection, operation, and analysis.

Role concentration and its risks.

A scope map showing the narrow anchor of one model on one product and the boundaries on generalization.

Model and product scope.

1.8.2 Documentation census of the documentation experiment

The control contains more in-code documentation than the documented build, while the stripped twin removes that layer by construction. The comment-line and docstring census is therefore a construct-validity warning: the headline contrast compares a curated trace layer with ordinary commenting rather than documentation with its absence.

No committed campaign census owns a comparable comment-line total for the three builds. The figure therefore makes no numerical density claim.

source: repository search of 5. Experiment/ for a committed cross-build comment census, 2026-09-14

The control’s comment-line mismatch and its history are recorded in the control-build history record.

1.8.3 Run scatter of the documentation experiment

The paired response matrix records the token delta for each cell. Run-to-run scatter shows how much the observed response varies across repetitions and therefore limits the stability of treatment comparisons.

source: 6. Metrics/paired_response_matrix.csv, column valid_for_primary_analysis

Seven run sets have a validity flag that reads true in the registry. They form the basis for the primary analysis.

source: 5. Run Sets/run_index.csv, column valid_for_primary_analysis

A scatter display showing paired response variation across valid runs and the distance between treatment responses.

Run-to-run scatter and signal separation.

A ceiling diagram showing the remaining uncertainty after a flawless run and the limits of the resulting claim.

The ceiling on a flawless confirmatory run.

1.8.4 Record conflicts of the documentation experiment

The repository contains governing records that disagree about validity, scope, and status. The record-validity figures are carried without a source by declaration until the responsible record is found. This conflict affects which observations can support a primary claim.

A reconciliation diagram comparing record-validity flags with the validity status used for analysis.

Record-validity flag reconciliation.

The specification count also shifts between governing records. The discrepancy is a traceability risk because a count can appear stable while its source changes.

A comparison diagram showing specification-count drift between governing records.

Specification-count drift.

1.8.5 Researcher roles of the documentation experiment

The researcher’s six roles create a connected threat across task selection, execution, analysis, and decisions. The protocol names the roles and the safeguards that separate their judgments where the records permit separation.

source: 5. Experiment/0. Plan/Appendix E, Protocol.md, section 4

1.8.6 Decision sheet confounds of the documentation experiment

The run-set timing confound is tracked when ordering or timing can differ across treatment conditions. The control-build model confound is tracked when the control’s model context can differ from the documented build. The Operator Decision Sheet assigns the run-set timing confound to decision item B9, the randomized dose-ladder entry, and the control-build model confound to decision item B10, the rebuild or observational-use entry.

source: 5. Experiment/0. Plan/Appendix I, Operator Decision Sheet.md, line 94 (item B9) and line 100 (item B10)

B9 keeps the dose ladder as a randomized interleaved block. B10 records whether the control is rebuilt or demoted to observational use. The control comment-line mismatch is linked to the control-build history record, and the task-selection history is linked to the retired removal suite record. source: 5. Experiment/0. Plan/Appendix I, Operator Decision Sheet.md, line 94 (item B9) and line 100 (item B10)

1.8.7 Open question status of the documentation experiment

Appendix H records open, resolved, and deprioritized questions. Those statuses mark which validity threats still require a decision and which have an adopted disposition.

source: 5. Experiment/0. Plan/Appendix H, Open Questions.md, line 5 (open), line 26 (resolved), and line 63 (deprioritized)

The weaknesses are recorded in the weaknesses entry.

Previous: Findings of the shakedown batches | Next: Pending freezes and registrations
source: operations/site-ia/A6-page-briefs.md, section "### R7. Threats to validity"