Chapter 1.1 opens part 1, the reading path, and the section as a whole; the numbered contents at its end lead into every later chapter, beginning with chapter 1.2, The claim and its falsification conditions.
The experiment asks whether documentation written into a codebase changes the reliability and token cost of maintenance work by an artificial-intelligence coding agent. The registered question concerns the effect of that documentation under defined task and analysis conditions. No confirmatory result, a result from the pre-registered analysis, exists yet. Every result measured so far belongs to the shakedown phase.
source: the folder 5. Experiment/0. Plan/
A source line, a citation identifying the repository location from which a claim or number is taken, follows the relevant claim. The site uses source lines for factual claims and keeps exploratory readings separate from confirmatory findings.
source: 0. The Website/NORTH_STAR.md lines 20 to 26
1.1.1 Experimental arms and measurement
A harness, a controlled tool for building and running the software, produces the experimental trees. An arm, one treatment version measured against the same tasks as the other treatment versions, is identified by an arm code and a tree name. The three arms are H-LAP-L5P, the harness-built tree using the Literate/Anchored Programming protocol (LAP) at the highest documentation dose; H-STR, the harness-built tree with its documentation mechanically stripped; and H-NON, the tree produced without the harness. Evidenceline, the product under study, supplies the codebase shared by these trees.
source: operations/site-ia/A6-page-briefs.md section ”# R0.”
A cell, one execution of one task against one arm under one seed, is the unit of measurement. A run set, a batch of cells executed together, groups work for monitoring. A token, one unit of text read or written by the agent, measures part of the work cost. The same maintenance tasks are measured against each arm.
A shakedown, a test of the measurement apparatus rather than a test of the treatment, is the status of all measured work so far. A confirmatory result is a result produced after the hypotheses, task set, and analysis rules have been registered and applied to eligible cells. No such result exists yet.
source: the folder 5. Experiment/0. Plan/
1.1.2 Current status of the apparatus
The live counts change with every completed batch review. The following status applies on the date of the last update, and the changelog row records that update date.
The experiment has 1,200 completed cells, including 1,058 valid cells and 142 invalid cells.
source: 5. Experiment/7. Monitoring/LIVE_SUMMARY.md line 5
The run index contains 274 run-set rows. Its largest identifier is 353; identifiers are assigned when runs are created, so the largest identifier need not equal the number of rows.
source: 5. Experiment/5. Run Sets/run_index.csv (274 data rows, last row 351-20260917-artifact-activation-v2-s10-neutral)
Project H has sixteen registered task units and one draft among the seventeen Project H rows.
source: 5. Experiment/4. Task Library/task_index.csv column status
The library contains eight documentation doses. The count uses profile identifiers in the specified rows and excludes repeated identifiers carrying the frozen status.
source: 5. Experiment/3. LAP Profile Library/lap_profile_index.csv rows 2 to 9 (rows 10 and 11 repeat L4P and L5P with the status frozen, so the drafter counts profile identifiers and not rows)
The Evidenceline product has thirteen trees in its trees folder.
source: the folder listing of 5. Experiment/2. Project Library/Project H - Evidenceline Compliance Platform/variants/
Amendment A15, ratified on 2026-09-20, closed the first phase of the experiment, in which agents added features to the purpose-built Evidenceline product, and opened a second phase in which agents repair defects deliberately seeded into imported open-source projects. The first imported project is changedetection.io, a web-page change monitor of 38,697 source lines, held as Project J in three arms: the untouched snapshot, a matched-information arm carrying orientation files and inline markers, and a full arm that adds the relationship map. Two seeded defects, J-M001 and J-M002, passed their baseline checks on all three arms; a third, J-M003, was set aside because the project’s own visible tests detect it. The calibration batch for the two admitted defects, run set 354, twenty cells on the untouched arm at two ticket levels, started on 2026-09-20 and has no results yet.
source: 5. Experiment/0. Plan/Appendix A (Chronology of Design Amendments) entry A15; 5. Experiment/2. Project Library/Project J - changedetection.io/source_manifest.json (size_at_import); 5. Experiment/4. Task Library/Project J/*/baseline_failure_check.md; 5. Experiment/5. Run Sets/354-20260920-brownfield-J-calibration-NON/
1.1.3 Reading path contents of the documentation experiment
The generated contents table records the entry titles for the experiment reading path. The generator writes it between these markers and regenerates it from the entry files.
The next entry is The claim and its falsification conditions. The deep-dive entries are Execution of a cell, The codebases and their variant trees, The profile library of documentation doses, and The maintenance tasks and their hidden checks. Terms are collected in Glossary of campaign terms.
source: the entry files in operations/site-ia/drafts/
1.1.4 Site map and status chart
The first figure presents the shared specification, the three arms, the common task set, and the measured outcomes.
Figure D-R0-1. The specification, the three arms, the dose ladder, the tasks, and the measured cell
%% figure D-R0-1 flowchart TB subgraph g1[" "] n1_1["5. Experiment/2. Project Library/Project H - Evidenceline Compliance Platform/2. Requirements/"] --> n1_2["variants/"] n1_2 --> n1_3["3. LAP Profile Library/lap_profile_index.csv"] n1_3 --> n1_4["4. Task Library/task_index.csv"] end subgraph g2[" "] n2_1["6. Metrics/cell_factor_matrix.csv"] --> n2_2["6. Metrics/statistical_analysis_plan.md"] end n1_4 --> n2_1
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section ”# R0.”
The experiment structure connects one frozen specification to three measured arms.
The site map is regenerated from the tree so that its entry names remain current.
Figure D-R0-2. The reading path, the four working folders, and the reference part
%% figure D-R0-2 flowchart LR subgraph g1[" "] n1["gen_site_figures.py"] end
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section ”# R0.”
The site map shows the reading path, deep dives, and glossary.
The status chart combines the pass-count matrix, the run index, and the live summary.
source: 5. Experiment/6. Metrics/pass_count_matrix.csv, 5. Experiment/5. Run Sets/run_index.csv, and 5. Experiment/7. Monitoring/LIVE_SUMMARY.md
Figure D-R0-3. The count of recorded cells beside the count that may enter the analysis
%% figure D-R0-3 flowchart LR subgraph g1[" "] n1_1["5. Experiment/7. Monitoring/LIVE_SUMMARY.md"] --> n1_2["valid_for_primary_analysis == true"] n1_2 --> n1_3["status == complete"] end
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section ”# R0.”
The status chart shows the current measurement status and run-set progress.
Contents of the section
Every entry of the section, numbered the way the plan is numbered. Part 1, the reading path, is written to be read in order from the first chapter to the last; parts 2 to 5 are deep dives into the harness, the project library, the profile library and the task library, and part 6 is the reference part. Each may be read on its own after the reading path.
Part 1. The reading path
- 1.2 The claim and its falsification conditions
- 1.3 The measuring instrument
- 1.4 The eight documentation doses
- 1.5 Measurement of one cell
- 1.6 The pre-registered analysis plan
- 1.7 Findings of the shakedown batches
- 1.8 Threats to validity
- 1.9 Pending freezes and registrations
Part 2. The harness
- 2.1 Execution of a cell
- 2.2 The cell pipeline from staging to scoring
- 2.3 Isolation by the container
- 2.4 The maintenance prompt and its versions
- 2.5 The guards of the harness
- 2.6 The token ledger and the timing record
- 2.7 The closeout chain of a batch
- 2.8 The model-serving routes
- 2.9 The configuration files and their defaults
- 2.10 The harness test suite
- 2.11 The future harness topology
Part 3. The project library
- 3.1 The codebases and their variant trees
- 3.2 The Evidenceline product
- 3.3 The project manifest
- 3.4 Differences between the variant trees
- 3.5 Construction and freezing of the builds
- 3.6 The control build and its history
- 3.7 The second codebase, paperless-ngx
- 3.8 The archived projects
Part 4. The profile library
- 4.1 The profile library of documentation doses
- 4.2 The six documentation artifacts
- 4.3 The purpose and rules of each dose
- 4.4 Physical differences between the rungs
- 4.5 Measurement of the dose
- 4.6 Application of a profile to a build
- 4.7 Weaknesses and measured uptake of the treatment
Part 5. The task library
- 5.1 The maintenance tasks and their hidden checks
- 5.2 Construction of a maintenance task
- 5.3 The retired seeded-defect design
- 5.4 Anatomy of one task package
- 5.5 Purpose of the hidden tier
- 5.6 Scoring of a cell
- 5.7 The chained task sequences
- 5.8 The witness and the verification records
- 5.9 Task applicability by arm
- 5.10 The retired removal suite and the draft tasks for other projects
Part 6. The reference part
- 6.1 Glossary of campaign terms
- 6.2 The campaign boundaries and their pooling rules
- 6.3 The batch ledger
- 6.4 The primary records and their order of precedence
- 6.5 The detailed design of the harness
Part 7. The detailed design of the harness
The fourteen chapters of the detailed design, one per mechanism of the harness, are published beside this section under the design folder; the entry 6.5 The detailed design of the harness introduces them for a general reader.
Part 8. The task deep dives
One page per maintenance task, built from the failure forensics, with a source line under every number: the task deep-dive index.