source: cell_forensics.csv; rollup.json
Task registration facts
The accession is H-M020, the identifier assigned to this maintenance task, and the host-side record of the task names it a ceiling on the duration of a single time entry. The task index records the task as registered for measured runs. An arm, one software build supplied to an agent, ran as H-LAP-L1, H-LAP-L1H, H-LAP-L2C, H-LAP-L2S, H-LAP-L2T, H-LAP-L3B, H-LAP-L4P, H-LAP-L5P, H-NON, H-STR. A cell, one execution for an arm and seed, is the row counted below. A seed, the repeat number that distinguishes repetitions, is part of the cell identifier. A run set, one named batch of cells, is part of the evidence record. A probe, a hidden check used to score work, can appear in a run-set identifier. A documentation dose, the amount of documentation supplied with a build, can appear in a run-set identifier. An agent-coder, a local model alias used for a coding run, can appear in a run-set identifier (arm identifiers).
source: 5. Experiment/4. Task Library/Project H/H-M020/task_manifest.json; 5. Experiment/4. Task Library/Project H/H-M020/README.md; cell_forensics.csv
Attempt outcomes by arm
A hidden tier, host-side checks withheld from the agent, decides the hidden result. A visible tier, checks shown during the attempt, can disagree with it. An invalid attempt, an attempt carrying one or more instrument flags, remains visible in its outcome count.
| Arm | Attempts | Hidden-tier pass | Hidden-tier fail | Not run | Invalid |
|---|---|---|---|---|---|
H-LAP-L1 | 12 | 9 | 1 | 2 | 3 |
H-LAP-L1H | 11 | 7 | 2 | 2 | 3 |
H-LAP-L2C | 1 | 1 | 0 | 0 | 0 |
H-LAP-L2S | 1 | 1 | 0 | 0 | 1 |
H-LAP-L2T | 1 | 1 | 0 | 0 | 0 |
H-LAP-L3B | 1 | 1 | 0 | 0 | 1 |
H-LAP-L4P | 1 | 1 | 0 | 0 | 0 |
H-LAP-L5P | 12 | 7 | 1 | 4 | 7 |
H-NON | 9 | 7 | 2 | 0 | 1 |
H-STR | 9 | 7 | 2 | 0 | 1 |
| source: cell_forensics.csv |
Recorded failure causes
| Failure cause | Count | Share of task rows | Example cells |
|---|---|---|---|
A2, a cause code (a grouped failure label) for success after a hidden-failure retry | 2 | 3.4% of 58 task rows | 102-20260910-auto-igx-thor-agent-coder-0210 / H-LAP-L5P__H-M020__full__s3; 132-20260911-auto-igx-thor-agent-coder-0308 / H-STR__H-M020__full__s6 |
A5, a cause code (a grouped failure label) for no edit made | 8 | 13.8% of 58 task rows | 095-20260909-auto-dgx-qwen3-4b-fp8-2125 / H-STR__H-M020__full__s3; 096-20260909-auto-igx-thor-qwen3-4b-fp8-2145 / H-NON__H-M020__full__s3; 114-20260910-auto-igx-thor-qwen3-4b-fp8-1535 / H-LAP-L1H__H-M020__full__s4 |
| source: rollup.json; cell_forensics.csv |
Visible and hidden tier disagreements
| Run set | Cell | Field | Expected record | Observed record |
|---|---|---|---|---|
203-20260912-auto-igx-thor-agent-coder-1053 | H-LAP-L5P__H-M020__full__s9 | population_key | register | missing_from_forensics (record comparison) |
| source: disagreements.csv |
Agent pass timing figures
The pass evidence contains no duration field, so it provides no median or longest agent pass figure. source: pass_forensics.csv
Evidence sources
The numbers above come from the named forensics files and the task manifest. The forensics package records provenance commit 6a0e5432f8750522e3bbacaf89b48de13a8e8058.
source: cell_forensics.csv; pass_forensics.csv; disagreements.csv; rollup.json; 5. Experiment/4. Task Library/Project H/H-M020/task_manifest.json; provenance.json