Chapter 5.2 follows chapter 5.1, The maintenance tasks and their hidden checks, and precedes chapter 5.3, The retired seeded-defect design, within part 5, the task library.
The customer ticket is the work order for a maintenance task. It states the requested behaviour and its acceptance criteria in customer language. Task families come from the product’s working use. A designer creates each task and an adversarial reviewer tests its specification. Admission is controlled before the task enters the measured suite, and its state is recorded for registration. Pre-run measures describe task difficulty.
5.2.1 Customer ticket work order
An additive task, a task that adds requested behaviour to an existing build without a prepared code change, gives each arm the same ticket and asks for behaviour to be added to its build. The ticket describes the work without planting a patch in the starting tree. This keeps the task independent of file layout and architecture. The acceptance criteria state what the completed behaviour must do, including observable outcomes and relevant boundary conditions.
The customer ticket provides a common work order without depending on a patch matching a target tree.
source: 5. Experiment/4. Task Library/
5.2.2 Task family structure of the maintenance-task suite
Task families are drawn from the product’s working use. An independent task is assessed on its own. A chained task carries the resulting tree from one step to the next, so the sequence records where the work stops meeting the ticket’s acceptance criteria. Both forms use the same customer-ticket model and the same admission requirements.
The task suite replaces removal-based work with additive work. The retired seeded-defect design remains an archived reference for its migration history. The task suite plan records the current family structure, the unit of analysis, and the transition away from the former design.
source: 5. Experiment/0. Plan/6. The Task Suite.md, task suite section
5.2.3 Design and review roles
The task designer grounds the requested behaviour in real code paths, checks the starting behaviour, writes the acceptance criteria, and prepares the hidden checks. The adversarial reviewer compares the task with the product specification and looks for wording that permits an unintended solution. The reviewer also checks that the task can be attempted against each eligible build without a planted change.
The designer and reviewer have separate responsibilities. The designer supplies evidence for the requested change. The reviewer searches for specification gaps, accidental clues, and acceptance checks that test a different behaviour. Their prompts define these responsibilities.
source: 5. Experiment/4. Task Library/
5.2.4 Admission decision of the maintenance-task suite
A realism gate, a decision point that admits only work resembling actual maintenance, checks the ticket before the task enters the measured suite. It compares the requested behaviour with the specification, the available builds, and the parity audit. A task is held for revision when the behaviour is already required, absent from every build for a shared reason, or duplicated by another task.
The gate preserves the reason for each rejection in the task library archive. Its review asks whether the task measures a maintenance change that a customer or maintainer could plausibly request. The gate is the basis for the realism review question and for the plan’s acceptance check.
Figure task_realism_gate. A ticket passes through specification, build, and parity reviews, with rejected tasks routed to an archive and an admitted task moving into the measured suite.
%% figure task_realism_gate flowchart LR n1["A ticket passes<br/>specification<br/>build<br/>parity reviews"] n2["rejected tasks routed to an archive and an admitted task moving into the measured suite"] n1 --> n2
source: operations/site-ia/A3-diagram-specification.md
The realism gate separates maintenance work from specification gaps, shared build shortfalls, and duplicate coverage.
source: 5. Experiment/4. Task Library/ and 5. Experiment/0. Plan/6. The Task Suite.md, realism gate section
5.2.5 Task record of the maintenance-task suite
A registration manifest, a record of a task’s identity, state, checks, and run eligibility, accompanies each task. It records the task type, difficulty band, seeded-defect state, hidden scorer, documentation obligation, status, and eligibility for measured runs. The manifest schema defines the allowed fields and values. The registration template supplies the record used during task preparation.
The manifest separates task description from admission state. Its status records whether preparation remains unfinished or a reason for redesign has been retained. The run-eligibility field is the value read by the coverage analysis. A task is eligible only after its checks have been traced to the ticket, exercised against the untouched builds, and shown to fail for the stated behavioural reason.
The registration template and schema are the controlling files for this record. The plan’s task-suite sections describe its place in the measured suite. The build and instrumentation method acceptance check covers registration and admission evidence.
source: 5. Experiment/4. Task Library/ and 5. Experiment/4. Task Library/task_manifest_schema.md
5.2.6 Difficulty profile of the maintenance-task suite
A difficulty factor, a pre-run measure of task difficulty, describes one aspect of the work that may affect an agent’s performance. The factor profile contains these names:
prompt_vagueness
source: 5. Experiment/4. Task Library/task_scoring_rubric.md line 5
source: 5. Experiment/4. Task Library/task_index.csv header row
invariant_density
source: 5. Experiment/4. Task Library/task_scoring_rubric.md line 5
source: 5. Experiment/4. Task Library/task_index.csv header row
search_directness
source: 5. Experiment/4. Task Library/task_scoring_rubric.md line 5
source: 5. Experiment/4. Task Library/task_index.csv header row
blast_radius
source: 5. Experiment/4. Task Library/task_scoring_rubric.md line 5
source: 5. Experiment/4. Task Library/task_index.csv header row
hidden_acceptance_pressure
source: 5. Experiment/4. Task Library/task_scoring_rubric.md line 5
source: 5. Experiment/4. Task Library/task_index.csv header row
repair_pass_risk
source: 5. Experiment/4. Task Library/task_scoring_rubric.md line 5
source: 5. Experiment/4. Task Library/task_index.csv header row
human_realism
source: 5. Experiment/4. Task Library/task_scoring_rubric.md line 5
source: 5. Experiment/4. Task Library/task_index.csv header row
The designer scores every factor before a run. The scores describe prompt clarity, non-local constraints, code-search distance, change spread, hidden acceptance demands, risk in the first repair, and the plausibility of the work for a human maintainer. A task with low human realism is excluded unless it serves as an explicit control.
5.2.7 Task suite transition of the maintenance-task suite
The task suite plan records the additive task family, its unit of analysis, and its departure from the retired seeded-defect design. The migration joins the former task-design material with the task-family material so that the ticket, review roles, realism gate, manifest, and factor profile form one construction path. The plan also records the destination of retired tasks.
The owning files are the task designer prompt, task reviewer prompt, realism gate, registration manifest template, task manifest schema, task scoring rubric, and task suite plan. Their acceptance evidence covers the plan path, the build and instrumentation method path, and the realism gate review.