The execution log and provenance record is the harness account of what happened while a run set was being made. A run set is the numbered directory containing one batch execution and its cell records. Provenance is the identifying information that says which requested model, served model, endpoint, prompt choice, workspace and process produced an observation. The mechanism is not the score of a cell and it is not the token or timing measurement. It is the connective record that lets an operator relate a decision by the batch driver, an invocation of the cell runner, an agent process, a timing span and an automatic scheduling decision to a file and a time.

The owning implementation is distributed across 5. Experiment/1. Harness/scripts/run_batch.py, 5. Experiment/1. Harness/scripts/run_cell.py, 5. Experiment/1. Harness/scripts/timing_lib.py, and 5. Experiment/10. User Interface/auto_experiment.py. The batch driver is the program that validates and iterates over a batch specification. The cell runner is the program that stages and runs one measured cell. A measured cell is one task, build condition, phase and repetition seed treated as one experimental unit. The timing recorder appends measured activity spans. Auto experiment mode is the server-side scheduler that chooses and launches later batches from coverage and operator limits.

The design writes a human-readable execution_log.md for run-set events, a machine-readable command_log.jsonl for agent invocations, per-cell timings.jsonl files for activity spans, and the scheduler files auto_experiment_state.json and auto_experiment_log.jsonl. It also writes provenance into cell_manifest.json, the per-cell manifest that records the requested configuration and the Fabric pass records. The records are consumed by the results interface, report and aggregation programs, and the operator. The source digest and the flow model identify the execution-log node as execution_log; they do not provide a separately labelled chapter 13 subgraph, so this chapter uses the model records and witnesses that name this node and its connected provenance records rather than inventing node identifiers.

1. The purpose and the position in the life of the experiment

The mechanism answers a plain question: what did the experiment attempt, what did it decline to attempt, what process ran, and which machine and model supplied the work? The batch driver invokes the cell runner once for each planned cell and may invoke it again when a hidden score failed while passes or budget remain. The cell runner invokes an agent backend and the scorer around each pass. The timing recorder is called by the runner and by the surrounding measurement code. The interface invokes the auto experiment tick on a timer, while auto_experiment.py itself calls functions supplied by the interface context to inspect a running unit, read a batch outcome, write a draft and launch a specification.

The record is produced incrementally. run_cell.py appends REACHED when a chained step is already satisfied by its start tree and appends NOT-RUN when a cell is refused or deliberately skipped. The batch driver appends NOT-RUN when a chain refusal is decided before the cell child is started. A deviation is appended when the parent repository changes during a guarded pass. The cell runner appends a command record after each agent invocation and writes the requested backend, redacted argument vector, pass, times and return code. The timing module appends one JSON object when a timed activity closes, and may maintain a live current-activity file while the cell is active. Auto experiment mode appends a scheduler event and rewrites its complete state on a change.

The records occupy different positions in the experiment life. The run-set execution log begins when a cell is reached or refused and remains useful through closeout. The command log is produced inside the pass loop. The cell manifest is first written after staging and the staging gate, then rewritten on Fabric lifecycle events and finalisation. Timing rows are written during execution and summarised after the cell. The scheduler state is read and saved on each tick, while its JSON-lines diary records the history of launches, holds and completions. This is therefore an operational audit trail, not a final report. A report may copy or interpret it, but the mechanism keeps the causal details close to the process that observed them.

The design does not claim that every file is immutable. execution_log.md and auto_experiment_log.jsonl are append-oriented, while auto_experiment_state.json, manifests and timing live markers are rewritten or replaced. The verified implementation gives the execution log a consistent append format and the state writer an atomic replacement operation. It does not establish a cryptographic append-only guarantee, and the distinction matters when an audit asks what was preserved rather than merely what the latest state says.

2. The reader’s map of the owning files

2.1 Batch orchestration

5. Experiment/1. Harness/scripts/run_batch.py is the batch orchestrator. Its constants and imports occupy lines 1 to 88, including the closeout chain and backend sets. Argument handling and signal setup are around lines 90 to 150. The core orchestration occupies lines 151 to 780. _log_not_run is at lines 227 to 241 and writes the run-set execution log when a cell is deliberately not run. _write_closeout_progress is at lines 580 to 620, and set_run_index_status is at lines 790 to 831 in the verified source despite the older digest grouping it with the core helpers. _closeout begins at line 833 and _closeout_locked at line 852. The public run_batch flow is described by the digest at lines 790 to 880, and the command-line main entry point is at the end of the file. Finalisation writes active-batch status, the batch outcome, run-index status, closeout progress and the closeout artifacts.

The important ownership boundary is that run_batch.py owns the batch-level decision to leave a cell unrun and the closeout status, but run_cell.py owns the normal reached and not-run line format. The batch helper deliberately uses the same wording as the cell helper so the interface does not need to know which process made the decision. The driver also invokes sibling programs, including run_cell.py and score_cell.py, so command and outcome records have a process boundary even when the batch operation is presented as one action.

2.2 Cell execution

5. Experiment/1. Harness/scripts/run_cell.py parses its command line in _parse_args at lines 1508 to 1549 and enters the main cell flow in run at line 1674. Setup, staging, manifest creation and timer construction occupy the early part of run; the pass loop begins at line 2362. The command record is appended at lines 2636 to 2640. The runner’s finalisation helpers are _log_reached at lines 3122 to 3132, _log_not_run at lines 3135 to 3139, and _finalize_manifests at lines 3142 onward. The signal conversion and executable entry point occupy lines 3207 to 3265.

The runner owns the per-cell provenance. It records the model and backend choice, prompt variant, retry-feedback mode, task and start tree, staging gate, billing and container settings, and, for a Fabric backend, the endpoint and pass lifecycle. It also owns the command log append made after an agent invocation. Its manifest callback rewrites the manifest when Fabric events arrive, so the latest admission, lease, model provenance and heartbeat state can be viewed while the pass remains active.

2.3 Timing recording

5. Experiment/1. Harness/scripts/timing_lib.py defines the timer at lines 88 to 277. StageTimer.__init__ is at lines 100 to 116, _append at lines 118 to 126, stage at lines 147 to 185, open at lines 187 to 203, and close at lines 205 to 229. The no-op timer is at lines 291 to 332. Reading begins with read_timings at line 337 and aggregation with summarize at line 352. Transcript and turn-ledger analysis follows later in the file. The file is below the digest threshold that would require a numbered comment-step inventory.

The timer’s finalisation action is the append on successful or exceptional exit from a stage context, followed by live-marker cleanup when the cell ends. A timing write failure is intentionally suppressed so loss of a diagnostic row does not turn a measured cell into a failed cell. The measured field distinguishes the harness stopwatch from an agent-reported or transcript-derived duration.

2.4 Automatic scheduling

5. Experiment/10. User Interface/auto_experiment.py keeps its constants and policy near lines 74 to 114. State loading is load_state at lines 210 to 236. Atomic state writing is save_state at lines 284 to 293. append_log is at lines 296 to 300 and log_event at lines 303 to 312. Decision functions are decide_next_batch at line 1478, decide_parallel_batches at line 1571, and build_spec at line 1729. The serial tick begins at line 1827 and the parallel tick begins at line 2072. The source file’s own comments define a tick as one action, such as hold, launch or completion observation, before saving and returning.

The scheduler’s state record is broader than the run-set execution log. It retains mode, settings, current units, endpoint failures, credit flags, closeout backfill warnings, daily launches, the last decision, a hold and a bounded history. auto_experiment_log.jsonl receives one JSON object per event. save_state writes a temporary file and replaces the target, so readers see either the old complete state or the new complete state. This is atomic replacement, not append-only history.

3. The inputs

3.1 Command-line arguments

The batch parser is in run_batch.py main, whose verified parser region is at lines 890 to 920 in the digest. It accepts --batch, --skip-preflight, --max-usd, --resume-into and --run-set-destination. The cell parser is run_cell.py _parse_args at lines 1508 to 1549. It accepts --run-set, --project, --condition, --task, --phase, --seed, --model, --prompt-variant, --retry-feedback, --agent-backend, --max-passes, --token-budget, --dry-run, --resume-hidden-fail, --base-dir, --fabric-endpoint and --start-tree. Auto experiment mode has no independent command-line parser. The server supplies a context to tick and the scheduler reads auto_experiment_state.json.

Program and parserArgumentDecision made
run_batch.py main, lines 890 to 920--batchBatch specification to load.
run_batch.py main, lines 890 to 920--skip-preflightWhether preflight is bypassed.
run_batch.py main, lines 890 to 920--max-usdCost gate limit.
run_batch.py main, lines 890 to 920--resume-intoExisting run set to continue.
run_batch.py main, lines 890 to 920--run-set-destinationExplicit destination for a new run set.
run_cell.py _parse_args, lines 1508 to 1549Cell identity and execution optionsRun set, project, condition, task, phase, seed, model, backend, pass limit and token budget.
run_cell.py _parse_args, lines 1508 to 1549Prompt and resume optionsPrompt variant, retry feedback, dry run and hidden-failure resume.
run_cell.py _parse_args, lines 1508 to 1549--fabric-endpoint and --start-treeFabric selection and chained start tree.

3.2 Environment variables

The verified digest lists the Fabric, billing and container variables as inputs to the cell runner. The execution-log mechanism does not read all of them itself. They decide the provenance that the runner places in the manifest and command context.

VariableRead or used atDecision
FIVEB_CELL_CONTAINERrun_cell.py source input region and backend setupContainer image or host execution route.
FIVEB_BILLINGrun_cell.py source input regionBilling route for the selected backend.
FIVEB_CELL_CREDENTIALS and ANTHROPIC_API_KEYrun_cell.py backend validationWhether the selected authentication route may start.
FIVEB_FABRIC_ENDPOINTrun_cell.py Fabric setup and dgx_fabric.py endpoint resolutionNamed Fabric endpoint when no explicit argument overrides it.
DGX_SPARK_FABRIC_MODELrun_cell.py model resolutionDefault local model alias.
DGX_SPARK_FABRIC_LANErun_cell.py Fabric setupAdmission lane.
DGX_SPARK_FABRIC_URL and DGX_SPARK_FABRIC_TOKENFabric endpoint resolution used by the backendController address and bearer credential. The token is not written as provenance.
FIVEB_LEASE_RETRY_SECONDSagent_backends.py lease retry setupInterval between retryable lease refusals.

3.3 Files read

File or pathReader and line or regionDecision
Batch specification named by --batchrun_batch.py main and run_batchPlanned cells, model, conditions and batch identity.
4. Task Library and task manifestsrun_batch.py and run_cell.py setupTask folder, chain definitions and visible checks.
Run-set template under 5. Run Sets/001-*run_batch.py _create_run_setInitial directory and template records.
config/run_defaults.json, config/model_matrix.jsonrun_cell.py setupDefaults, model and pass limits.
container/image_digest.txtrun_cell.py container identity helperExact image identity recorded with the cell.
Prompt templates and artifact hint filesrun_cell.py prompt setupPrompt version and optional prompt factor.
Existing cell manifest and earlier metricsrun_cell.py resume pathPrior configuration and pass state.
auto_experiment_state.jsonauto_experiment.py load_state, lines 210 to 236Scheduler mode, settings, current work and policy memory.
Run-set batch_outcome.jsonContext function called by tick, lines 1882 to 1893Status, reason and attempted-cell count after a batch.
closeout_progress.jsonauto_experiment.py _discover_closeout_backfill, lines 239 to 272Failed publication stages that need a warning or backfill.

3.4 Network endpoints

The batch driver and cell runner do not themselves add a network endpoint to the execution log. A selected backend may call the Anthropic API or a Fabric controller. Auto experiment mode receives live observations through the context assembled by server.py, including unit status, batch outcome and model catalog operations. The endpoint identity that matters for provenance is the selected backend and Fabric endpoint in the cell manifest. Secret tokens are inputs to the call, not record fields.

Client or context functionEndpoint or external operationProvenance consequence
Selected agent backendAnthropic service or Fabric inference serviceProvider and model identity in the pass result and manifest.
Fabric client used by the backendFabric controller admission and lease operationsEndpoint, admission, lease, heartbeat and release events in fabric.passes.<k>.
ctx.unit_stateLive launched unit statusHold while active or completion handling.
ctx.batch_outcomeBatch driver’s outcome recordScheduler event with status and reason.
ctx.start_specBatch launch operationLaunch event and current-unit state.

4. The happy path in order

4.1 Batch specification and run-set preparation

The batch path begins in run_batch.py main at lines 890 to 920, which parses the batch arguments and calls run_batch. run_batch begins at line 790 and validates the specification, selects or creates a run set, and enters the cell loop. _create_run_set is the allocation helper at lines 150 to 180. The design reason stated by the source digest is to allocate a unique numbered directory and copy the template, so records from one execution do not merge with another. The driver marks the active batch as running, then carries the planned cell identity into each child invocation.

The driver invokes _run_one_cell at lines 330 to 430. It calls the cell runner and scorer as separate child operations. If the first score has a hidden failure and passes or budget remain, the driver sets the hidden-failure resume flag and repeats the cell and score operations. The reason for this arrangement is that a visible success with a hidden failure can receive another measured pass without changing the cell folder. Each child result remains distinguishable through its cell and pass fields, while the batch retains the overall order.

4.2 Cell staging and identity

The cell path begins in run_cell.py run, lines 1674 to 3029, after _parse_args has read the required cell identity at lines 1508 to 1549. The runner resolves the task and variant, stages the working copy, applies the seeded defect or feature-removal diff, and runs the staging gate. It then creates the private workspace boundary, snapshots the staged tree and computes hashes. The design reason in the source description is to ensure that the agent receives the intended start tree and that the hidden checks do not already pass unless the step is recorded as reached.

The runner writes the cell manifest after staging and the gate. The manifest is the canonical provenance record for the cell because it joins requested factors with the realised start tree, prompt settings, container identity, billing route and Fabric details. A step already satisfied by its start tree uses reached_step_metrics, then _log_reached at lines 3122 to 3132. It is called REACHED rather than NOT-RUN because the step has a folder and records and has been established without an agent pass. A refusal uses _log_not_run at lines 3135 to 3139, naming the refusal reason.

4.3 Pass timing and command capture

The pass loop begins at line 2362. The runner constructs a StageTimer for the cell and opens spans around setup, agent invocation, transcript processing, visible testing, scoring and finalisation. StageTimer.stage begins at lines 147 to 185 in timing_lib.py; its exit path appends a timing row even when the enclosed operation raises. The reason stated by the timing module is diagnostic isolation: a missing timing line must not make the measured cell fail.

The backend is invoked inside the pass loop. After the process ends, the runner appends a command object through _append_jsonl at lines 2636 to 2640. The object carries run_id, cell_id, pass, an argument vector with the prompt replaced by <prompt>, prompt length, backend, start and end timestamps and the process return code. Replacing the prompt contents prevents the command log from becoming a second copy of the agent conversation while preserving the invocation shape needed for audit.

4.4 Agent result and cell continuation

The runner reads the handoff and transcript-derived usage, runs visible checks and, when those pass, invokes hidden scoring. It generates the pass patch and checks parent-repository fingerprints and invocation errors. The pass loop stops when the visible and hidden conditions pass, when the token budget is exhausted, when the maximum pass count is reached, or when a refusal or invalidity guard halts the cell. A hidden failure with remaining capacity returns to the loop with the locked retry notice. The design reason is to distinguish an allowed measured retry from an unbounded retry and to preserve the failure cause in the pass records.

At the end of run, _finalize_manifests begins at line 3142 and updates the run manifest and campaign run index. The run-set closeout is then reached through run_batch.py _closeout at line 833 and _closeout_locked at line 852. The closeout writes active status, batch_outcome.json, run-index status and initial closeout progress before invoking each closeout script. It rewrites progress around each invocation so a reader can see pending, running, finished, skipped or failed stages. The source comment gives the design reason: a failure in one report or aggregation program must not prevent later closeout stages from running.

4.5 Scheduler observation and launch

The scheduler path begins in auto_experiment.py tick, lines 1827 to 1905 and continuing through the serial decision path. It loads state with load_state, recovers expired endpoint failures and expires credit flags. If mode is off it saves and returns. If a current unit is active, it writes a hold with _set_hold and returns. If a unit has ended, it finds the run set, reads the batch outcome and calls log_event with a batch_finished entry.

When no batch is active, the scheduler validates settings, calls decide_next_batch at line 1478 or the parallel decision function at line 1571, builds a specification with build_spec at line 1729 and calls the supplied launch function. log_event at lines 303 to 312 appends the event to auto_experiment_log.jsonl and inserts the newest entry into the bounded state history. save_state at lines 284 to 293 then atomically replaces the state file. The design reason is that the scheduler should take at most one externally visible action per tick and leave enough state for the next tick or the browser to explain that action.

5. The state machine

The state values below are record-visible values, not a claim that the process has only these Python control states. The flow model identifies execution_log as the run-set text record, with REACHED, NOT-RUN and DEVIATION lines. It identifies command_log as the per-pass invocation record, the lease member of the cell manifest as the Fabric lifecycle record, and the timing records as stage observations. The following model subgraph is therefore rendered as Mermaid from the verified record and function witnesses for execution_log, command_log, lease, metrics, batch_outcome and the scheduler state. The node identifiers used here are batch.prepare, cell.reached, cell.not_run, cell.running, pass.invoked, pass.completed, pass.failed, cell.halted, batch.closed, scheduler.off, scheduler.hold, scheduler.launch and scheduler.finished.

stateDiagram-v2
    [*] --> batch_prepare
    state "batch.prepare" as batch_prepare
    state "cell.reached" as cell_reached
    state "cell.not_run" as cell_not_run
    state "cell.running" as cell_running
    state "pass.invoked" as pass_invoked
    state "pass.completed" as pass_completed
    state "pass.failed" as pass_failed
    state "cell.halted" as cell_halted
    state "batch.closed" as batch_closed
    state "scheduler.off" as scheduler_off
    state "scheduler.hold" as scheduler_hold
    state "scheduler.launch" as scheduler_launch
    state "scheduler.finished" as scheduler_finished
    batch_prepare --> cell_reached : run_cell 1933; execution_log REACHED
    batch_prepare --> cell_not_run : run_batch 227..241 or run_cell 3135..3139; NOT-RUN
    batch_prepare --> cell_running : run_cell 1674; manifest written
    cell_running --> pass_invoked : run_cell 2362..2640; command_log row
    pass_invoked --> pass_completed : returncode 0 and scoring continues
    pass_invoked --> pass_failed : nonzero return or guard
    pass_failed --> pass_invoked : hidden failure and budget remains
    pass_failed --> cell_halted : refusal or invalidity guard
    pass_completed --> cell_running : another pass permitted
    cell_running --> batch_closed : run_cell 3142 and final status
    cell_reached --> batch_closed : reached record finalised
    cell_not_run --> batch_closed : not-run record finalised
    batch_closed --> scheduler_finished : tick 1882..1893 reads outcome
    scheduler_finished --> scheduler_launch : decision and start_spec
    scheduler_launch --> scheduler_hold : unit_state active or validation refusal
    scheduler_hold --> scheduler_launch : later tick and eligibility restored
    scheduler_off --> scheduler_launch : operator turns mode on
    scheduler_launch --> scheduler_finished : launch event logged

The state machine intentionally separates REACHED from NOT-RUN. The source comment at run_cell.py lines 3122 to 3127 says that the interface and reports use not-run lines as the skipped-cell list, while a reached step has its own folder and records. It also separates pass.failed from cell.halted: a failed pass may be a normal retryable observation, while a halted cell carries an invalidity or refusal consequence. The scheduler states are state-file decisions rather than execution-log line values.

6. The sequence of one unit of work

A unit of work here is one measured cell pass, from the batch child invocation to the records that remain after the pass. Programs, the agent process, the Fabric service when selected, and the files are separate participants. The diagram is the hand-authored audit sequence required by the diagram plan, using cited functions and records rather than implying that the flow model supplied a chapter-labelled sequence.

sequenceDiagram
    participant Operator as operator or interface
    participant Batch as run_batch.py
    participant Cell as run_cell.py
    participant Timer as StageTimer
    participant Agent as agent process
    participant Fabric as Fabric controller or API service
    participant Manifest as cell_manifest.json
    participant Commands as command_log.jsonl
    participant Timings as timings.jsonl
    participant Log as execution_log.md
    participant Score as score_cell.py
    Operator->>Batch: launch batch specification
    Batch->>Cell: invoke one cell child
    Cell->>Manifest: write staged identity and configuration
    Cell->>Timer: open pass and agent spans
    Cell->>Agent: invoke backend with workspace and prompt
    Agent->>Fabric: request model service when local backend is selected
    Fabric-->>Agent: response and process outcome
    Agent-->>Cell: transcript, handoff and return code
    Cell->>Commands: append pass command record at run_cell 2636..2640
    Timer->>Timings: append closed timing rows
    Cell->>Score: score visible and hidden results
    Score-->>Cell: verdict and invalidity information
    Cell->>Manifest: finalise cell and pass provenance
    Cell->>Log: append REACHED, NOT-RUN or DEVIATION when applicable
    Cell-->>Batch: result JSON and exit status
    Batch->>Batch: decide next cell or hidden-failure resume
    Batch->>Log: append NOT-RUN for a driver refusal when applicable

The Fabric participant is conditional. A Claude route may use the Anthropic service, and a local route uses the Fabric admission, lease, heartbeat and inference lifecycle. The execution log itself is not a transcript of every response. The command log preserves the invocation envelope, the manifest preserves identity and lifecycle fields, and the timing file preserves measured spans. This separation avoids putting secrets or the full prompt into the human-readable event log.

7. The records

The canonical row for the run-set event log is the execution_log row owned by this chapter. It is written by run_cell.py _log_reached, _log_not_run and the parent-repository deviation helper, by run_batch.py _log_not_run, and by the interface stop path cited in the flow model. The canonical row for the command log is also here because the command is the provenance bridge between a process and a pass. The cell manifest and lease record remain canonical in the cell execution and Fabric chapters, while this chapter records how they join the audit trail. Every reader listed below is taken from the verified flow model or source grep. A reader listed as a file means that its reading is distributed through that file rather than concentrated in one function.

RecordWriter and lineFields or contentReaders
execution_log.mdrun_cell.py _log_reached, lines 3122 to 3132; _log_not_run, lines 3135 to 3139; parent-change helper at the source-digest witness; run_batch.py _log_not_run, lines 227 to 241; interface stop pathTimestamp, cell identifier, event word REACHED, NOT-RUN or DEVIATION, and reason5. Experiment/10. User Interface/server.py _not_run_entries, line 2856; gen_run_set_report.py; lint_report_prose.py
command_log.jsonlrun_cell.py _append_jsonl, called at lines 2636 to 2640run_id, cell_id, pass, redacted argv, prompt_chars, agent_backend, started, ended, returncodeaggregate_timings.py line 260 and capture_evidence.py line 119
cells/<cell>/cell_manifest.jsonrun_cell.py main flow after staging and on finalisation; Fabric callback rewrites it at the cited runner callbackSchema and run identity, cell factors, task and start tree, prompt settings, staging gate, container and billing details, Fabric endpoint and per-pass lifecyclerun_cell.py resume and manifest conflict checks; score_cell.py; aggregate_metrics.py; aggregate_timings.py; capture_evidence.py; gen_run_set_report.py; interface server
cell_manifest.json member fabric.passes.<k>_FabricLeaseBackend.run_pass, agent_backends.py lines 835 to 1061, through _record_fabric at run_cell.py lines 2522 to 2532Idempotency key, admission and lease identifiers, admission state and reason, exact model provenance, timestamps, wait seconds and rounds, agent and release durations, endpoint, context window and lifecycle eventsaggregate_timings.py; run_cell.py model-provenance finalisation; interface live view
cells/<cell>/timings.jsonltiming_lib.py StageTimer._append, lines 87 to 93 in the digest and implementation regionSchema, run and cell IDs, stage, phase, pass, parent, depth, start and end, seconds, ok, measured, attributestiming_lib.py read_timings and summarize; aggregate_timings.py; interface timing views
cells/<cell>/timing_current.jsontiming_lib.py StageTimer._mark_live, lines 96 to 113, overwritten during active workSchema, IDs, update time, current stage and phase, pass, start and active stackInterface live timing view
run_manifest.jsonrun_cell.py _finalize_manifests, lines 3142 onward, with record line in the flow model at 3184Run roster, cell status, not-run reason, visible result, budget state, Fabric endpoint and primary-analysis validitygen_run_set_report.py, gen_iteration_summary.py, cleanup_run_set.py, attest_batch_outcome.py, restore_reached_step.py and interface server
5. Run Sets/run_index.csvrun_cell.py finalisation under the publish lock and run_batch.py set_run_index_statusOne campaign row per run set, including status and experiment identityupdate_monitoring.py, interface server, publish_snapshot.py, and run_batch.py at the status update
batch_outcome.jsonrun_batch.py _write_batch_outcome, lines 617 to 648Schema, run ID, status, reason, cells planned, attempted, skipped and not reachedgen_run_set_report.py and attest_batch_outcome.py; interface reads it through its batch-outcome helper
closeout_progress.jsonrun_batch.py _write_closeout_progress, lines 685 to 711, and _closeout_locked lines 852 to 895Schema, run ID, batch status, overall status and times, and seven stages with key, title, script, status, times, duration, message and artifactsbackfill_closeout.py, interface server _closeout_progress, and auto_experiment.py backfill discovery
auto_experiment_state.jsonauto_experiment.py save_state, lines 284 to 293Mode, settings, session and iteration records, current units, lane outcomes, endpoint failures, credit flags, closeout backfill, daily launches, last decision, hold and bounded historyauto_experiment.py load_state and tick; interface server and browser auto panel
auto_experiment_log.jsonlauto_experiment.py append_log, lines 296 to 300, called by log_event lines 303 to 312UTC minute, event name and detail objectInterface history and operator inspection; the scheduler itself keeps the newest fifty entries in state
cells/<cell>/metrics.jsonrun_cell.py main finalisationPass totals, budget and validity, result fields and timing summary, including exact served-model identity where availablescore_cell.py, aggregate_metrics.py, interface views and reports
cells/<cell>/transcripts/pass<k>.stream.jsonl and handoff filesCell runner and agent backendAgent interaction stream and handoff note used to explain a command result and derive timing or usagetiming_lib.py, score_cell.py, audit_transcripts.py, reports and interface views

The table keeps a record with its writer even where this chapter is not the sole owner. A Fabric pass record is not a second execution log. It is the detailed provenance below the cell manifest, and the execution log can name a deviation or refusal without containing lease secrets. Similarly, batch_outcome.json is the canonical batch termination record, while execution_log.md supplies the chronological reason lines that explain individual decisions.

The lineage diagram shows how a process observation becomes a campaign observation. A campaign table is a derived table, meaning that it is produced from preserved run-set records rather than written by the pass itself. The diagram names the source files and the principal programs that read them.

flowchart TD
    Spec[batch specification] --> Batch[run_batch.py]
    Batch --> Outcome[batch_outcome.json]
    Batch --> Closeout[closeout_progress.json]
    Cell[run_cell.py] --> Manifest[cell_manifest.json]
    Cell --> Log[execution_log.md]
    Cell --> Commands[command_log.jsonl]
    Timer[timing_lib.py StageTimer] --> Timings[cells/<cell>/timings.jsonl]
    Backend[agent backend and Fabric] --> Manifest
    Agent[agent transcript] --> Metrics[metrics.json]
    Manifest --> Aggregate[aggregate_metrics.py]
    Metrics --> Aggregate
    Timings --> AggregateTiming[aggregate_timings.py]
    Commands --> AggregateTiming
    Log --> Report[gen_run_set_report.py]
    Outcome --> Report
    Closeout --> Interface[server.py and operator interface]
    Aggregate --> Campaign[6. Metrics campaign tables]
    AggregateTiming --> Campaign
    Campaign --> Monitoring[monitoring, reports and interface]
    State[auto_experiment_state.json] --> Scheduler[auto_experiment.py tick]
    Scheduler --> Diary[auto_experiment_log.jsonl]
    Scheduler --> Spec

The lineage is not a claim that every downstream program reads every upstream file. aggregate_timings.py reads command records to reconstruct missing timing spans and reads the per-cell timing and lease material. gen_run_set_report.py reads run-set records for a complete report. The campaign tables are outputs of aggregation, not authoritative replacements for the per-cell verdict or the event line. The interface may read a live manifest or scheduler state before a closeout table exists, which is why the state files and the append-oriented diaries remain separate.

8. The loops and the waits

8.1 Batch cell loop

run_batch.py iterates the cells in the specification in run_batch at lines 790 to 880. Its bound is the length of the planned cell list. It stops at the end of that list, when the batch is stopped, or when the consecutive agent-invocation-error rule aborts the remaining cells. It does not retry a normal cell automatically; the only immediate continuation is the explicitly allowed hidden-failure resume handled by _run_one_cell at lines 330 to 430.

8.2 Hidden-failure resume loop

The resume loop is bounded by the cell’s maximum pass count and cumulative token budget. It continues only when the hidden result failed and both the pass and budget rules permit another attempt. It stops after a successful visible and hidden result, at the maximum pass count, at budget exhaustion, or at a halt. The retry is immediate at the driver level. There is no stated sleep between the original scoring result and the next run_cell.py child.

8.3 Closeout loop

run_batch.py _closeout_locked iterates the fixed CLOSEOUT_CHAIN at lines 852 to 895. The bound is the number of configured closeout stages. The loop stops after the final stage, not at the first failed stage. Each child invocation is guarded, its status is written to closeout_progress.json, and the next stage is attempted. The retry policy is none in this function. A later operator or backfill tool can rerun a failed closeout stage, but that is outside this loop.

8.4 Cell pass loop

run_cell.py starts the pass loop at line 2362. The bound is max_passes together with the token budget. The stop conditions are visible and hidden success, a refusal, an invalidity halt, budget exhaustion, maximum passes, or a termination signal converted by the signal handler. A hidden-failure retry is immediate and receives the selected retry notice. An agent invocation error does not receive a normal pass retry inside the cell; it halts the cell and lets the batch count it.

8.5 Timing and transcript loops

timing_lib.py reads every JSON line in a timing file through read_timings and summarises every retained row through summarize. Its bound is the number of file lines or rows present. Transcript timing and turn-ledger construction loop over transcript events and tool calls. They stop at end of file. There is no wait and no retry; malformed or unreadable material is skipped or reduced to an empty or partial result according to the source digest.

8.6 Scheduler tick loop

The interface calls auto_experiment.py tick periodically. The verified interface documentation gives a thirty-second tick interval. One tick loads state, makes at most one hold, launch or completion observation, saves state and returns. The loop stops when the server or mode is stopped. It does not retry within the same tick. A failed launch is recorded and the next tick applies the normal waiting and eligibility rules rather than spinning in the current call.

8.7 Endpoint cooldown and credit expiry

The scheduler records a failed endpoint and applies ENDPOINT_FAILURE_COOLDOWN_MINUTES, whose source value is ten minutes at line 108. The endpoint is eligible again on a later tick after the cooldown expires. This is an automatic delayed retry, not a repeated request loop. Paid-route credit flags expire at the next UTC midnight, as described by expire_credit_flags at lines 315 onward. No interval or request retry is performed while the flag is active.

8.8 Fabric waits represented in provenance

The execution-log chapter does not own the Fabric admission loop, but the provenance record carries its wait result. The backend waits between retryable lease refusals using the configured retry interval and records waiting_for_lease events, wait rounds and wait seconds in the per-pass lease member. Fabric admission polling also waits between controller state reads using its polling interval. These waits have no execution-log line for every poll. Their evidence is the manifest lifecycle and the timing rows that the runner places around the pass.

9. The guards and refusals

Guard or checkLocationRefusal or altered flow
Batch preflight and cost gaterun_batch.py main path, lines 890 to 920Exit before run-set creation when preflight or the cost limit refuses the specification.
Duplicate cell identityrun_batch.py run_batch, lines 790 to 880Refuses the batch rather than allowing two planned cells to share an identity.
Mixed model batchrun_batch.py run_batch, lines 790 to 880Refuses a batch that names incompatible model choices under the batch contract.
Existing run-set destination or invalid resume pathrun_batch.py _create_run_set, lines 150 to 180, and resume setupRaises a system exit rather than merging records into an existing or incomplete directory.
Chain predecessor absent or refusedrun_batch.py _chain_start_tree and _note_chain_outcome, lines 230 to 310Does not run the dependent cell and writes a not-run reason.
Task or variant resolutionrun_cell.py setup within run, lines 1674 to 1933Refuses the cell when the task folder, variant root or realization diff is unavailable.
Contaminated task or artifact briefrun_cell.py scan in the setup pathRefuses before agent invocation and records the refusal reason.
Staging gaterun_cell.py staging-gate path before manifest creationRefuses when the hidden start-tree condition is not established, unless the step is the explicit reached case.
Cell manifest conflictrun_cell.py check_cell_manifest_conflict, cited by the digest at line 1584Refuses reuse under a changed model, prompt variant, retry mode or Fabric endpoint.
Missing start treerun_cell.py chained start-tree setupRefuses the cell rather than starting from an unproven tree.
Prompt lengthrun_cell.py prompt compositionTruncates the retry notice when the prompt cap would be exceeded and continues with a bounded notice.
Patch sizerun_cell.py patch creation and validationHalts the cell when the patch exceeds the configured cap.
Parent repository fingerprintrun_cell.py parent_repo_fingerprint and deviation helperHalts or marks the cell invalid when guarded parent files changed during the pass and appends DEVIATION.
Agent invocation resultrun_cell.py invocation-error guardHalts with agent_invocation_error and invalidity rather than treating a failed process as a scored attempt.
Reported usage unavailablerun_cell.py provider usage handlingHalts the local provider path as usage_unavailable and marks it invalid.
Environment-dependent failurerun_cell.py pass result logicHalts after the second consecutive environment-dependent failure.
First termination signalrun_cell.py _sigterm_to_systemexit, lines 3207 to 3230Converts the first signal to SystemExit(143) so finalisation can release records; later signals are ignored.
Scheduler modeauto_experiment.py tick, lines 1869 to 1871Saves and returns without launch when mode is off.
Running unitauto_experiment.py tick, lines 1873 to 1880Sets a batch_running hold and returns.
Settings validationScheduler decision path after the current-unit branchSets a validation hold and logs the reason instead of launching.
Daily launch, token and credit limitsScheduler policy constants and decision functionsHolds or excludes an otherwise eligible candidate.
Endpoint failure quarantineScheduler endpoint-failure recovery and decision pathExcludes the failed endpoint until cooldown expiry and records the failure.
Closeout publication failureauto_experiment.py backfill discovery, lines 239 to 272Keeps a closeout backfill warning in state and history rather than claiming complete publication.

The guards have two different evidentiary roles. A refusal before the manifest exists is explained by the execution log and child result, while a refusal after staging has a manifest and metrics record. A guard that alters rather than stops the flow, such as prompt truncation or endpoint cooldown, must be read with the changed field or scheduler state. The log line alone is not sufficient to reconstruct the full altered path.

10. Every unhappy path

The following table uses four parts for each path. The first part states what the step is supposed to do. The second states why the implementation works that way. The third states the trigger and the record left. The fourth states the cost and downstream effect. A downstream effect is stated for the driver, scorer, aggregator and interface even when the effect is that no program is reached.

Trigger and code pathIntended step and design reasonRecord, status and exit codeCost and downstream effect
Preflight or cost gate refusal in run_batch.py mainValidate the specification before allocating a run set, so an impossible or unaffordable batch leaves no partial roster.JSON error on standard output, no run-set execution log, and exit code 1.Costs validation time. The driver does not enter the cell loop, the scorer and aggregator receive no cell, and the interface can show the launcher error rather than a running batch.
Duplicate cell ID or mixed-model validation in run_batch.py run_batchReject an internally ambiguous batch, so one cell directory and one model contract have one meaning.System exit before a run set is created; no cell record.Costs specification validation. The driver stops before child invocation, the scorer and aggregator do nothing, and the interface has no completed run-set record to display.
Invalid resume path or existing destination in run-set creationResume only into a real run set or create only a new destination, so records are not merged accidentally.System exit naming the invalid path or existing destination. A partial run set is not claimed by this chapter as a new record.Costs path checks. No scorer or aggregator child is called; the interface sees the launcher failure or the prior run set unchanged.
Missing task folder, variant root or realization diff in run_cell.py setupStage the exact task and start tree before agent work, so a pass cannot be attributed to an absent or wrong input.SystemExit or refusal result with task folder not found, variant root not found, missing_defect_realization or missing_feature_realization; a NOT-RUN line is written where the refusal path reaches _log_not_run, and the cell status is not_run.Costs filesystem resolution. The batch may mark the cell skipped, the scorer has no valid agent pass, the aggregator excludes or classifies the not-run record, and the interface reads the reason from the log or manifest.
Contaminated brief or artifact hintKeep study and harness information out of the agent-facing task, so the task remains the assigned maintenance problem.Refusal reason contaminated_task_brief or contaminated_artifact_hint, with NOT-RUN and not-run status when finalisation is reached.Costs the scan. The driver can continue to later cells, the scorer does not score this cell, the aggregator records no measured attempt, and the interface shows a refused cell.
Staging gate failureEstablish that the hidden condition needs repair before spending agent time, so a pre-satisfied or unscorable task is not counted as an agent success.Refusal reason staging_gate_failed, staging output and not-run record when the cell reaches the logging helper. Status is not_run; the child returns the not-run exit convention used by main, which is 2 for a not-run result.Costs staging and hidden-check time. The driver records a skipped cell, the scorer is not meaningfully run, the aggregator sees a gate refusal, and the interface can display the gate detail.
Cell manifest conflictPrevent a second configuration from occupying a cell folder, so provenance remains one-to-one with the cell identity.System exit with the relevant model, prompt, retry-feedback or Fabric-endpoint conflict. The existing manifest remains the record.Costs manifest comparison. The driver records the child refusal and may mark the cell not run; the scorer and aggregator do not overwrite the existing cell, and the interface continues to show the original configuration.
Parent repository change during the passDetect an agent or external process changing guarded parent files, so the measured workspace boundary is not silently broken.DEVIATION line from the parent-change helper, halt status protocol_violation, invalid_for_primary_analysis true, and preserved metrics or manifest details. The child returns the halt result rather than a normal success.Costs fingerprint checks and the remaining cell opportunity. The driver counts a halted result, the scorer marks or propagates invalidity, the aggregator excludes it from primary analysis, and the interface shows a protocol violation.
Nonzero agent process or API invocation errorCapture the agent failure and stop the cell, so an unavailable process is not confused with a failed maintenance attempt.NOT-RUN line naming agent_invocation_error where the refusal logger is reached, status agent_invocation_error, invalidity true, command-log return code and stderr or transcript evidence. run_cell.py main returns the result-derived non-success code.Costs the attempted process and error extraction. The driver increments its consecutive error count and aborts remaining cells after the configured two consecutive errors; the scorer marks the cell invalid or cannot score it, the aggregator excludes it from primary results, and the interface reports the invocation error.
Usage unavailable on a local providerRequire usable usage evidence for a measured route, so a timing or performance row cannot be detached from its budget accounting.NOT-RUN reason usage_unavailable, invalidity true and provider record as far as it was written.Costs the pass and usage inspection. The driver receives a halted child, the scorer propagates invalidity, the aggregator excludes the primary measurement, and the interface shows the usage failure.
Two consecutive environment-dependent failuresStop a cell that cannot establish the required execution environment, so retries do not spend indefinitely on the same host problem.Halt result with environment_dependent_code and the corresponding per-pass evidence; a not-run line is emitted on the final halted path when that path calls the logger.Costs the failed passes. The driver treats the cell as halted, the scorer marks the invalid reason unless waived, the aggregator excludes it from primary analysis, and the interface shows the environmental cause.
Budget exhausted or maximum passes reachedEnd an otherwise valid attempt at the declared experimental bound, so the agent cannot run beyond the planned exposure.Metrics flags budget_exhausted or pass-limit completion, final manifest status and per-pass command and timing rows. There need not be a NOT-RUN line because the cell did run.Costs the declared budget or pass time. The driver proceeds to scoring and the next cell, the scorer records the final verdict, the aggregator uses the bounded outcome, and the interface shows attempts and budget state.
Termination signalRelease or finalise what can be released while representing that the batch stopped, so a stale running status is not left behind.First signal becomes SystemExit(143) in the cell; the batch closeout writes stopped_before_completion in batch_outcome.json and closeout progress.Costs cleanup and closeout time. The driver runs its finally closeout, the scorer and aggregators process completed records where possible, and the interface reads stopped status and closeout progress.
Scheduler settings validation failureCheck model, build, task, prompt, cap and endpoint settings before launch, so auto mode does not create a known-invalid batch.Hold object with validation detail, scheduler history entry and auto_experiment_log.jsonl event; no child exit code because no batch process is launched.Costs one tick. The driver, scorer and aggregator are not invoked, and the interface shows the hold and its reason.
Scheduler sees a running unitWait for the launched batch, including closeout, so two automatic launches do not overlap their shared campaign outputs.State hold with code batch_running; no new launch event.Costs the tick only. The existing driver, scorer and aggregator continue normally, and the interface shows the hold and current unit.
Scheduler launch failureAttempt one launch and retain the failure, so a disk, subprocess or endpoint problem cannot cause a tight retry loop.State hold and scheduler diary entry with failure detail; a credit flag is added when the reason matches the credit policy. There is no child batch exit code available when start_spec itself fails.Costs the launch attempt and the configured wait before another tick. No scorer or aggregator receives a new batch, and the interface shows the failed launch and hold.
Endpoint failure during unit or outcome inspectionQuarantine an endpoint that cannot answer, so an unhealthy local route is not selected repeatedly.Endpoint failure state with failure time and cooldown, plus a scheduler diary event.Costs the failed inspection and ten-minute cooldown. The driver is not started on the quarantined endpoint, the scorer and aggregator see no new cell, and the interface shows endpoint unavailability.
Closeout stage failure or publication backfillPreserve stage-level closeout progress, so one report or publication failure does not erase the metrics already produced.closeout_progress.json stage status failed, message and artifacts; scheduler state may carry a closeout backfill warning and log event. The batch outcome still records the batch status separately.Costs the failed stage and later backfill work. The driver continues later closeout stages, the scorer is already finished, aggregators retain successful outputs, and the interface shows the failed stage rather than claiming full publication.
Timing-file write, read or transcript parse failureKeep measurement diagnostics best effort, so an ancillary timing problem does not destroy the cell result.Missing or partial timings.jsonl, skipped malformed rows, empty timing summary or partial turn ledger; the cell continues without a new process exit code.Costs diagnostic precision only. The driver and scorer retain the cell result, aggregation marks what can be derived, and the interface may show missing timing detail.

11. The metrics produced or fed

The mechanism produces provenance and diagnostic fields first. Aggregators and reports decide which of those fields become campaign columns. A field is advisory when it explains timing, routing or validity without being a treatment outcome or a score.

Source field or recordDownstream column or useProgram and status
cell_id, condition, task, phase and seed in the manifestCell identity and factor columns in cell and campaign tablesaggregate_metrics.py copies the realised cell identity; these are analytic keys, not advisory.
Requested model, backend, billing route and Fabric endpointModel class, provider and endpoint fields in cell metrics and interface viewsrun_cell.py finalisation and aggregation use them to prevent silent route substitution. Endpoint is an explanatory factor.
Exact served-model provenance in fabric.passes.<k>Exact model identity in metrics.json and model grouping in reportsThe runner carries the last pass provenance into metrics. This is a provenance field and is essential when an alias can resolve to different served identities.
Prompt variant and retry-feedback modePrompt factor columns and report groupingThe runner records them in the manifest and metrics; the aggregator copies them into factor data. They are treatment or design factors when the batch uses them, not timing metrics.
execution_log.md event and reasonNot-run, reached and deviation display; report narrative evidenceserver.py parses not-run entries and gen_run_set_report.py reads the log. The text is advisory explanation unless a flow witness explicitly requires it for a state.
command_log.jsonl start and endReconstructed agent span and command timing inputsaggregate_timings.py uses the pass key and timestamps when a timing row is absent or needs placement. It feeds timing, not score.
command_log.jsonl backend, argv and return codeContainer evidence and invocation auditcapture_evidence.py extracts Docker invocation lines. This is audit evidence and does not become a performance score.
timings.jsonl stage, phase, pass and secondsStage totals, pass timing rows, cell timing rows and timing summary tablestiming_lib.py summarises locally and aggregate_timings.py writes campaign timing tables. Container stages are excluded from totals where the timing contract says they contain nested work.
timing_current.json current stage and stackLive activity label and elapsed-time displayThe interface reads it as advisory live status. It is overwritten and removed when the cell finishes.
lease_wait_seconds and lease_wait_roundsMachine-wait timing and explanation of queued workFabric lifecycle data feeds timing aggregation and live view. It is advisory to task success but material to wall-clock comparisons.
metrics.json pass result and validity flagsCell metrics, validity and primary-analysis inclusionscore_cell.py, aggregate_metrics.py, reports and interface consume it. Validity fields alter analysis inclusion.
batch_outcome.json status and countsRun-set status, stopped or aborted labels and report completenessThe report and attestation tools read it. It is the canonical batch outcome, not a per-cell score.
closeout_progress.json stage status and durationCloseout tracker, publication backfill and interface statusThe interface and backfill tool read it. It is advisory to the experiment result but authoritative for whether publication stages completed.
Scheduler event, hold and endpoint detailAuto history, hold reason and operator displayauto_experiment.py retains the latest fifty history entries and the diary retains all appended entries. This is operational provenance, not a campaign metric.

The route and model fields are the central provenance tags. A requested alias alone is not enough to identify a served model, so a local pass records the controller result in the manifest when the backend provides it. The command record deliberately replaces the prompt argument with <prompt>, which preserves invocation provenance without turning the command log into a transcript. Timing fields may be absent because StageTimer._append suppresses write errors. Such absence is a diagnostic limitation, not evidence that the activity took zero seconds.

12. The tests

The verified test inventory includes 5. Experiment/1. Harness/scripts/tests/test_run_batch.py, test_run_batch_sequences.py, test_run_cell.py, test_reached_step_scoring.py, test_resume_into.py, test_closeout_progress.py, test_aggregate_timings.py and test_timing_lib.py. The interface inventory includes 5. Experiment/10. User Interface/test_auto_experiment.py, test_live_cell_states.py, test_reports.py and test_timing_views.py. The source grep also verifies Fabric endpoint wiring tests that assert credentials do not enter command_log.jsonl.

test_run_cell.py exercises execution-log reached and not-run records, command-log redaction and cell finalisation paths. test_reached_step_scoring.py covers the reached status and its REACHED line. test_run_batch_sequences.py covers chain refusal and the resulting not-run entry. test_resume_into.py covers forwarding and resume argument construction. test_closeout_progress.py covers atomic progress writes, stage transitions and unwritable-progress behavior. test_aggregate_timings.py covers command-log reconstruction. test_timing_lib.py covers timing rows, live markers, clearing and suppressed diagnostic failure. test_run_batch.py covers batch iteration, retries, invocation errors and closeout orchestration.

test_auto_experiment.py covers off mode, active-unit holds, completion observation, settings and launch decisions, endpoint failure behavior, credit flags, history and closeout backfill. test_live_cell_states.py covers the interface interpretation of execution-log lines. test_reports.py covers report reading of run-set records. test_timing_views.py covers interface reading of timings.jsonl and omission of excluded stages. These tests exercise the mechanism through seams and temporary run sets rather than a live batch.

The mapping from every row in the unhappy-path table to one individual test was not made. The following coverage statement is therefore a test inventory, not a claim of one-to-one path coverage. Batch validation, resume, chain refusal, signal and closeout behavior are covered by the named batch and closeout tests. Cell staging refusals, reached records, command records, manifest conflicts and invocation errors are covered in test_run_cell.py and the reached and resume tests. Timing read and write degradation is covered by test_timing_lib.py and test_aggregate_timings.py. Scheduler holds, launch failures, endpoint failures and credit expiry are covered in test_auto_experiment.py. The exact combination of a parent-repository deviation, a downstream scorer result, an aggregator decision and an interface rendering is not identified by a single verified test here. The exact combination of every closeout publication failure and every downstream reader is likewise not mapped.

13. The dated incidents that shaped the code

The comments and nearby design notes carry several dated incidents that explain why the mechanism is shaped as it is. They are included as historical constraints, not as a list of defects in the draft.

On 2026-08-29, the harness changed termination handling so a signal is forwarded to the running child and the run records stopped_before_completion rather than leaving a stale running status. The batch README states that the closeout was moved into a finally path on 2026-08-31 after stopped run sets 041 and 050 lost their metric rows. This explains why _closeout is called from the finally path and why batch_outcome.json is written before the closeout children.

The same README records a 2026-08-31 repair to the run-index status. Cells had been writing active as they finished, but the final status was not being rewritten, leaving historical rows describing completed batches as active. The batch driver’s set_run_index_status and closeout order are the provenance consequence: a cell-level status is provisional until the batch closes.

On 2026-09-01, the closeout chain gained a written analysis stage after fifteen consecutive run sets had complete numbers but no recorded interpretation. This incident shapes closeout_progress.json and the scheduler’s publication backfill. A failed analysis or publication stage is a named stage failure, not a reason to discard the metric rows that came before it.

On 2026-09-03, the harness README records the addition of the turn ledger. It preserves one line per model request and tool call beside the pass transcript, and the timing module reconstructs those rows after the agent process exits. This explains why the command log remains a compact invocation record while detailed request provenance belongs to pass<k>.turns.jsonl.

On 2026-09-05, the auto experiment module addendum added the request preset to the model coverage identity and added a session iteration record. The code comments define a preset as a bundle of request settings pinned for a lease and explain why two otherwise similar model runs must remain separate coverage slots. The same dated addendum explains why the scheduler keeps a bounded iteration record in state while retaining its full event diary.

The timing module comments also record a general operational decision: diagnostic writes and reads are best effort. The reason is that a disk-full or read-only timing path must lose a timing line rather than lose a measured cell. The execution-log and command-log paths are therefore not interchangeable with timing diagnostics. The former explain control decisions and invocations; the latter measure spans and may be incomplete without making the cell invalid.

14. The weakest claim, what was not checked, and the token line

The weakest claim is that the combined records provide a complete audit trail. The source confirms the writers, formats, named readers and many refusal witnesses, but it does not establish that append-oriented files cannot be edited after the process exits, that every downstream reader has been exhaustively traced, or that every unhappy path has a dedicated test. The chapter also did not verify a separately labelled chapter 13 subgraph in the flow model because the inspected model uses record keys and witnesses rather than a chapter 13 label. The diagrams identify the record and function nodes used here and mark the sequence and lineage diagrams as hand-authored where the diagram plan requires it.

Model: openai/gpt-5.6-luna via OpenRouter; tokens: see the job ledger