1. The purpose and the position in the life of the experiment

This mechanism gives an operator a browser view of experiment evidence and gives the experiment an optional autonomous way to choose and start the next batch. A run set is a folder containing the cells and result artefacts for one batch. A cell is one planned programming attempt identified by condition, task, phase, seed and repetition. A batch is the group of cells sent to the batch driver together. The interface is the FastAPI service in 5. Experiment/10. User Interface/server.py; the automatic loop is its background decision service, implemented in auto_experiment.py.

The operator invokes the interface by opening the static application or by sending its HTTP requests. The server reads the campaign matrix, run-set records, monitoring records, cell manifests, controller status and service status. It returns JSON for the browser and writes only the records associated with drafting, tagging, control actions and automatic scheduling. It does not execute a cell itself. It asks a systemd user service, the operating-system supervisor used here for a background process, to run the batch launcher. The driver, scorer and aggregators then create the measured records that the interface displays.

The automatic loop is invoked by the server startup task and, through the auto controls, by the operator. It wakes after thirty seconds, takes a lock shared with requests that change automatic state, and runs one tick. A tick loads persistent state, removes expired failure and credit restrictions, observes any current unit, decides whether work is eligible, writes a draft specification, and sends that specification through the same _start_spec path as a manual launch. Its output is a state file, an append-only event log, and, when a launch is accepted, a systemd unit and a batch specification. Its ordinary purpose is to keep eligible resources working without silently bypassing model, endpoint, wait, credit and daily-count restrictions.

The live page presents five externally visible batch states: idle, starting, running, stopped_resumable, and aborted. stopped_resumable means that a batch was stopped before completion and still has a run set that the operator may resume. aborted is the terminal operator decision that prevents that resume. Automatic mode has four values in its persistent record: off, on, finishing, and hold. finishing means that the current batch may finish but no next batch is to be launched. hold is represented by the hold reason field while the mode remains active and the loop waits for a condition to change.

The position in the life of an experiment is therefore after planning and before execution for a launch, and alongside execution for observation. Manual drafting reaches preflight and the driver through the interface. Automatic drafting reaches the same launch path through a context supplied by the server. During execution the live endpoint reads process, run-set and cell records. After execution the next tick reads the batch outcome and closeout progress, updates the iteration record, and either chooses another batch or records a refusal to proceed. PhD publication is a separate control surface exposed by the same server and is not part of the cell execution path.

2. The reader’s map of the owning files

The principal owner is 5. Experiment/10. User Interface/server.py. Its verified regions are the read endpoints cells at lines 100 to 110 and runsets at lines 113 to 118, coverage and catalogue handling at lines 230 to 350, specification drafting, validation and saving at lines 360 to 510, and live state assembly in live at lines 2117 to 2160 and its continuation through the endpoint body. Control and launch handling is distributed across _stop_batch_and_release at lines 2320 to 2369, control_pause at lines 2372 to 2378, control_stop at lines 2381 to 2437, _launch_unit at lines 2440 to 2458, control_resume at lines 2562 to 2580, _start_resource_pool_spec at lines 2592 to 2674, _start_spec at lines 2677 to 2785, and control_start at lines 2788 to 2799. Automatic settings and state endpoints are at lines 4309 to 4410. The automatic background wiring is _auto_tick_locked at lines 4448 to 4452, _auto_tick_loop at lines 4455 to 4484, static serving at lines 4486 to 4517, the idle watchdog at lines 4520 to 4528, and main at lines 4531 to 4561. The file comments use dated numbered operational changes rather than a single numbered step list. The comments around _start_spec explain static preflight, delayed version-control work and the 2026-08-31 failure it prevents. The comments around _auto_tick_loop explain why an exception is logged without killing the loop.

5. Experiment/10. User Interface/auto_experiment.py is the automatic decision owner. The state region includes default_state at lines 181 to 207, load_state at lines 210 to 236, _discover_closeout_backfill at lines 239 to 272, save_state at lines 284 to 294, append_log at lines 296 to 301, and log_event at lines 303 to 312. The decision and specification helpers occupy the middle of the file, including settings validation, slot construction, coverage selection, decide_next_batch, decide_parallel_batches, and build_spec; their public entry points are verified in the code digest. The numbered comments identify the tick as section 6.2 at lines 1812 to 1818. The main flow is tick at lines 1827 to 2069 and tick_parallel at lines 2072 onward. Finalisation is the branch that fills session.iterations and clears current at lines 1882 to 1950, plus the parallel lane finalisation at lines 2099 to 2140 and its continuation. The file’s own comments call out the state file, the one-action-per-tick rule, session iteration records, endpoint isolation, credit flags and closeout backfill.

5. Experiment/10. User Interface/scheduler_lp.py is the immutable scheduling model used by the interface’s scheduling support. Lines 1 to 37 contain imports and definitions preceding the records. The frozen records include Job, Worker, Assignment, AggregateRecord and PublishRecord in the first 134 lines. The Scheduler flow is plan at lines 135 to 153, complete at lines 155 to 163, aggregate at lines 165 to 173, publish at lines 175 to 182, report_failure at lines 184 to 197, probe at lines 199 to 202, and idle_reason at lines 204 to 233. Helpers are at lines 243 to 287. This file is a scheduling value model, not the writer of the automatic state file.

5. Experiment/10. User Interface/research_state.py supplies the research-state view. Its public functions are verified in the source inventory as load_state, save_state, merge_settings, summary and default_state; the file is the state transformation support used by the interface rather than the automatic launch owner. It reads and normalises research records and returns the summary consumed by the API. No separate persistent record owned by this chapter is asserted for it without a verified writer path.

5. Experiment/10. User Interface/phd_publish.py owns the publication subprocess boundary. Its entry point is main at lines 205 onward. The server exposes publication status and push handlers at lines 3428 to 3433. The browser module app/phd-publish.js calls GET /api/phd-publish/status and POST /api/phd-publish, refreshes every 60 seconds while visible, and delays a post-push refresh by 1.5 seconds. Publication is adjacent to the results interface, but it is not a transition in the automatic batch state machine.

The static shell is 5. Experiment/10. User Interface/app/index.html. It provides the DOM containers, navigation, stylesheets, Plotly and the application modules. The browser modules include app.js, research-state.js, auto.js, timing.js, summary.js, reports.js, evidence.js, cohorts.js, phd-publish.js, and the other view modules listed by the digest. They render server responses and transient browser state. start_ui.sh, stop_ui.sh and publish_ui.sh are operational wrappers. The interface README and the two local specifications explain operation, but this chapter treats source code as the authority for a function, record or line range.

3. The inputs

3.1 Command-line arguments

ArgumentParser lineDecision made
--hostserver.py 4532 to 4534Address on which uvicorn binds, default 127.0.0.1.
--portserver.py 4533 to 4535Port on which uvicorn binds, default 8850.
--idle-minutesserver.py 4535 to 4538Whether the idle watchdog is enabled and its timeout.

The server parser does not define the draft’s claimed --base, --run-sets, --app-dir, --lanes or --auto-loop-enabled arguments, so they are not inputs to this verified design. The browser supplies request JSON instead. control_start reads path and optional skip_preflight at lines 2789 to 2799. control_resume reads run_set and optional skip_preflight at lines 2563 to 2577. Automatic settings arrive through the settings endpoint at lines 4309 onward.

3.2 Environment variables

VariableVerified readWhat it decides
Server path and configuration variablesmodule configuration before the endpointsThe resolved base, run-set, application and launcher paths used by the imported server. The exact variable names are not retained here because a verified source line was not located in the reviewed range.
Fabric endpoint and credential variables_fabric_client and its Fabric client construction, used by _stop_batch_and_release at lines 2360 to 2366Controller address and credentials for a lease release.
Automatic loop flagserver.py 4436 to 4452_AUTO_LOOP_ENABLED, which is a source constant and startup guard, not an environment variable.

The server source verified for this chapter reads configuration through imported path and Fabric helpers. It does not support the draft’s invented line 3935 environment-variable table. The server’s command line and request bodies are therefore the stable inputs stated here. Secrets are passed to the Fabric client and are not returned in the records described below.

3.3 Files read

File or familyVerified readerDecision or view supplied
6. Metrics/cell_factor_matrix.csvcells and coverage helpers, lines 100 to 110 and 230 to 270Planned cells, factors and coverage gaps.
5. Run Sets/run_index.csvrunsets, lines 113 to 118Run-set index shown to the operator.
7. Monitoring/monitor_state.jsonrunsets, lines 113 to 118Monitoring state.
5. Run Sets/run_tags.jsontag readers used by _start_spec, lines 2771 to 2780Existing tags and automatic tag membership.
5. Run Sets/cohort_presets.jsoncohort API readersComparison presets.
Task library filescoverage and draft helpers, lines 230 to 270 and 509 onwardTask and chain definitions used for a draft.
5. Run Sets/<run set>/ artefactslive lines 2117 to 2160 and report helpersCell state, outcomes, closeout and reports.
5. Run Sets/active_batch.md_parse_active and live helpersCurrent or paused batch marker.
5. Experiment/10. User Interface/auto_experiment_state.jsonauto_experiment.load_state lines 210 to 236Automatic mode, settings, current work, restrictions and history.
5. Experiment/10. User Interface/auto_experiment_log.jsonlauto_experiment.append_log writes it; the server context exposes the automatic log to the interfaceAppend-only automatic events.
systemd unit state and journal_unit_failure_for lines 2506 to 2553Whether an accepted launch died before producing a run set and its exit status.

The automatic tick also receives matrix rows, chains, live tasks, process state, active-batch state, unit state, batch outcome and closeout progress through ctx. These are server callables, not direct network inputs to auto_experiment.py.

3.4 Network endpoints

EndpointClient functionPurpose
GET /api/livebrowser fetch from the live viewCurrent batch state, processes, controller and cell information.
POST /api/control/startcontrol_startValidate and launch a supplied specification.
POST /api/control/pausecontrol_pauseStop the current work while retaining resumability where the driver permits it.
POST /api/control/stopcontrol_stopStop and abort selected run sets and disable automatic mode.
POST /api/control/resumecontrol_resumeLaunch a resume of a paused run set.
Fabric controller API_fabric_client, called at server.py 2360 to 2366Release a lease during operator stop.
systemd user manager_launch_unit, lines 2440 to 2458Start the batch driver.
systemd user manager and journal_unit_failure_for, lines 2532 to 2547Inspect the accepted unit and retrieve failure text.
GET /api/phd-publish/status and POST /api/phd-publishapp/phd-publish.jsInspect and request publication.

4. The happy path in order

4.1 Server startup

main in 5. Experiment/10. User Interface/server.py lines 4531 to 4561 parses host, port and idle timeout, registers the startup handler, optionally creates the idle watchdog, creates _auto_tick_loop, and starts uvicorn. The code comment gives the design reason for creating the automatic task only here: importing the module or constructing a FastAPI test client must not launch a real batch. The static root at lines 4486 to 4489 serves index.html, and RevalidatingStatic at lines 4492 to 4517 adds Cache-Control: no-cache so a browser checks changed scripts rather than showing an old dashboard.

4.2 Browser assembly and observation

The browser loads app/index.html, its styles and modules. The application modules attach to the DOM and fetch API data. The live module requests /api/live; live begins at line 2117, derives active processes and the run set, and assigns the visible state at lines 2126 to 2139. A process with a run set is running, a process without one is starting, a paused active marker is stopped_resumable, and no such evidence is idle. The response also contains the controller view, generation time and the selected run-set details. The comments give the reason for preferring a process that names its run set over the shared active marker: concurrent endpoint lanes can rewrite that marker independently.

4.3 Draft creation and validation

The operator submits a draft to the specification endpoint at lines 509 onward. The server fills coverage gaps and applies task and chain rules. The save endpoint at lines 667 onward writes a draft specification and returns a launch hint. This is a plan, not an execution record. The operator can then send its path to control_start at lines 2788 to 2799. The handler resolves the path below the base and passes it to _start_spec, keeping HTTP formatting outside the launch body.

4.4 Static preflight and commitment

_start_spec in server.py lines 2677 to 2785 first refuses a path that is not an existing file under the run-set directory, then refuses an unreadable specification. Unless the request skips it, it calls the static preflight at lines 2722 to 2732. The code comment gives the reason: a plan that cannot run should be reported to the person pressing Launch instead of failing minutes later in a background log. A non-numbered draft is copied to a numbered committed specification at lines 2734 to 2740. A tag is attached to the predicted run set at lines 2758 to 2781 after a successful launch request. The prediction is visible if the background unit never creates the folder.

4.5 Systemd launch

_launch_unit in server.py lines 2440 to 2458 constructs a unit name, invokes systemd-run --user with the batch launcher and its arguments, and returns either an error or launched: True with the unit. The subprocess has a 60 second timeout. The code comment distinguishes acceptance of the command from survival of the batch, because a service may pass this boundary and later fail preflight. _start_spec returns the unit, specification and optional tag to either the human endpoint or the automatic context.

4.6 Automatic tick entry

_auto_tick_loop in server.py lines 4455 to 4484 waits for _AUTO_TICK_SECONDS, which is 30 at line 4437, and runs _auto_tick_locked in an executor. _auto_tick_locked lines 4448 to 4452 holds _AUTO_LOCK while it calls auto_experiment.tick through _auto_context. The code comments give both reasons: file and subprocess work must not block the ASGI loop, and a request that changes the same state file must not interleave with a tick. An exception is appended as tick_exception and does not end the loop.

4.7 State recovery and completion handling

tick in auto_experiment.py lines 1827 to 2069 begins by calling load_state at lines 1840 to 1842. load_state lines 210 to 236 treats missing, empty, damaged or non-object JSON as default_state, repairs an old session without iterations, and discovers unresolved publication failures. The tick expires endpoint failures and credit flags, selects serial or parallel operation, and returns immediately with saved state when mode is off at lines 1866 to 1871. If current names a unit, it checks the unit. An active unit causes a batch_running hold at lines 1873 to 1880. A finished unit is matched to its run set and outcome, its open session iteration is closed, a paid-route credit failure may be flagged, and current is cleared at lines 1882 to 1949.

4.8 Eligibility and decision

When no current batch remains, the tick refuses to compete with an unrelated running process at lines 1952 to 1956 and holds when the last batch is paused and resumable at lines 1958 to 1964. finishing becomes off between batches at lines 1966 to 1972. The configured wait interval is enforced at lines 1974 to 1986. validate_settings is called at lines 1988 to 1992. The context then supplies matrix rows, chains and live tasks, and decide_next_batch is called at lines 1994 to 1999. A negative decision becomes a recorded hold at lines 2001 to 2005. This is the point where coverage, enabled models, selected tasks, prompt variants, retry feedback, endpoint availability, caps and previous history affect the next action.

4.9 Specification construction and shared launch

A positive decision is converted by build_spec and written by ctx.write_draft at lines 2007 to 2009. ctx.start_spec invokes the server’s _start_spec, so automatic and manual launches share path checks, static preflight, numbering, tags and systemd. On success, the tick writes current with the unit, label, specification path, predicted run set, launch time, backend, model, request preset, prompt variant, retry feedback and cells at lines 2011 to 2029. It increments daily and session counts, appends an open iteration, clears hold, and logs batch_launched at lines 2030 to 2058. On failure it records launch_failed, sets hold, and sets the last-finished time so repeated failures cannot spin at lines 2059 to 2067. save_state then completes the tick.

4.10 Operator stop and resume

control_pause at lines 2372 to 2378 sends the blocking _stop_batch_and_release work to an executor. That helper terminates units or processes, waits for death, and releases an outstanding Fabric lease at lines 2360 to 2366. control_stop at lines 2381 to 2437 appends an abort line to each targeted run set’s execution_log.md, replaces a paused active marker with aborted_by_operator, and changes automatic mode to off, clearing current work and logging auto_off. control_resume at lines 2562 to 2580 finds the committed specification and starts it with --resume-into. Thus pause preserves a resumable state where the driver permits it, while stop is the explicit terminal action.

5. The state machine

The flow model 5. Experiment/11. Detailed Design/flow-model/flow_model.v001.json does not model this chapter: its open item open-24 lists chapters 01, 02, 07, 11, 12 and 13 as outside the first stage, and it carries no node with a chapter value of 11. The first diagram below is therefore hand-authored from server.py and auto_experiment.py, with the line ranges named on each transition, and its node identifiers (prefixed c11_) are proposed for the next version of the flow model rather than read from it. The proposed identifiers are c11_ui_launch, c11_auto_start, c11_auto_tick, c11_auto_tick_execute, c11_auto_core_tick, c11_auto_decide, c11_auto_build_spec, c11_auto_launch, c11_auto_hold, c11_auto_credit_flag, c11_auto_endpoint_failure, c11_auto_closeout_backfill, c11_auto_save_state, c11_auto_log_event, c11_auto_off and c11_ui_stop. The model supplies the automatic-loop transitions and record witnesses. The live batch states are hand-authored from live and control_stop, because the plan explicitly calls for the interface view and the model seed does not contain those live nodes.

stateDiagram-v2
    [*] --> idle
    idle --> starting: control_start, server.py 2788-2799
    idle --> starting: control_resume, server.py 2562-2580
    starting --> running: live sees process and run set, server.py 2126-2139
    starting --> idle: unit failure, server.py 2145-2153
    running --> stopped_resumable: pause completes, server.py 2372-2378
    running --> aborted: control_stop, server.py 2381-2437
    stopped_resumable --> running: control_resume, server.py 2562-2580
    stopped_resumable --> aborted: control_stop, server.py 2387-2417
    aborted --> [*]

The live state is an observation derived from process and file evidence, not a separately maintained enum. starting lasts while a process exists without a run-set identity. running requires both. The stopped and aborted values are read from the active marker and the abort record.

The next diagram is also the chapter 11 model subgraph rendered as Mermaid, with the same node identifiers. It names the record-writing nodes that the JSON model places on failure and completion edges. The state labels off, on, finishing and hold are the persistent automatic-loop values. hold is a reason-bearing condition in the state record rather than a distinct mode value in the verified tick code.

stateDiagram-v2
    [*] --> off
    off --> on: c11_auto_start, auto_start endpoint
    on --> on: c11_auto_tick -> c11_auto_tick_execute
    on --> hold: c11_auto_decide, no eligible slots
    hold --> on: later tick finds eligible decision
    on --> on: c11_auto_build_spec -> c11_auto_launch
    on --> finishing: c11_auto_off, auto_off endpoint
    finishing --> off: c11_auto_core_tick, server.py 1943-1949
    on --> off: c11_ui_stop, server.py 2423-2437
    on --> hold: c11_auto_launch, launch_failed
    on --> on: c11_auto_credit_flag, credit flag recorded
    on --> on: c11_auto_endpoint_failure, endpoint failure recorded
    on --> on: c11_auto_closeout_backfill, closeout warning recorded
    c11_auto_tick --> c11_auto_tick_execute
    c11_auto_tick_execute --> c11_auto_core_tick
    c11_auto_core_tick --> c11_auto_decide
    c11_auto_decide --> c11_auto_build_spec: decision ok
    c11_auto_build_spec --> c11_auto_launch
    c11_auto_launch --> c11_auto_save_state
    c11_auto_save_state --> c11_auto_log_event
    c11_auto_log_event --> c11_auto_tick: next 30 second interval

A state transition is not inferred merely from a browser colour. The mode field, current, currents, hold, session.ended_reason, lane_outcomes, credit_flags, endpoint_failures and closeout_backfill are the durable witnesses. A failure may therefore leave the loop active and on hold rather than turning it off.

6. The sequence of one unit of work

A unit of work here is one automatic decision that becomes one supervised batch launch. The diagram is a hand-authored structural diagram from the verified elements in server.py and auto_experiment.py, not a rendering of the flow model. It includes the browser, server event loop, automatic worker thread, state files, systemd and the batch driver. The Fabric controller is shown because the stop path can release its lease, although lease acquisition belongs to the batch driver and its harness chapters.

sequenceDiagram
    participant Browser as Browser
    participant Server as server.py
    participant Loop as asyncio auto loop
    participant Worker as executor thread
    participant State as auto_experiment_state.json
    participant Diary as auto_experiment_log.jsonl
    participant Systemd as systemd user manager
    participant Driver as batch driver
    participant Fabric as Fabric controller

    Browser->>Server: GET /api/live
    Server->>State: load automatic state through auto context
    State-->>Server: mode, current, hold, history
    Server-->>Browser: live JSON and rendered view
    Loop->>Loop: sleep 30 seconds
    Loop->>Worker: run _auto_tick_locked
    Worker->>State: load_state
    State-->>Worker: current state or default
    Worker->>Worker: validate, observe, decide_next_batch
    Worker->>State: write_draft via server context
    Worker->>Server: start_spec(draft_path)
    Server->>Systemd: systemd-run --user batch unit
    Systemd-->>Server: accepted or nonzero result
    Server-->>Worker: launched, unit, spec, tag or error
    alt launch accepted
        Worker->>State: set current and session iteration
        Worker->>Diary: append batch_launched
        Worker->>State: atomic save_state
        Systemd->>Driver: start batch driver
        Driver->>Fabric: acquire and use lease as required
    else launch refused
        Worker->>State: set hold and last_batch_finished_at
        Worker->>Diary: append launch_failed
        Worker->>State: atomic save_state
    end

The server thread and browser path are separated from the automatic worker by _AUTO_LOCK. The lock covers the state read, decision and write for one tick, but not the browser’s rendering. The tick delegates process and subprocess calls to its context. A unit accepted by systemd is not a completed unit of experiment work. The next tick must observe the service and its run-set outcome before it closes the iteration record.

7. The records

The following table gives the canonical records owned by this chapter. A record is named with its writer and every reader that was verified or is spread through a named reader file. The cell manifests, scores, timing tables and campaign aggregates are written by other chapters. This chapter reads them and therefore does not duplicate their canonical rows.

Record or tableWriter and lineFields or contentsReaders
5. Experiment/10. User Interface/auto_experiment_state.jsonauto_experiment.py, save_state lines 284 to 294schema_version, mode, settings, session, current, currents, lane_outcomes, lane_finished_at, endpoint_failures, last_batch_finished_at, credit_flags, closeout_backfill, daily_launches, last_decision, hold, historyauto_experiment.py load_state lines 210 to 236; server.py automatic context and auto endpoints; app/auto.js through the automatic API.
5. Experiment/10. User Interface/auto_experiment_log.jsonlauto_experiment.py, append_log lines 296 to 301, called by log_event lines 303 to 312One JSON object per line with at, event and detail; the flow model calls these timestamp, event and detailsserver.py automatic context and API support; app/auto.js History view; operators through the file.
Draft specification JSONserver.py automatic context write_draft, and the draft endpointLabel, cells and selected model, backend, prompt, retry, endpoint and request policy values produced by build_spec_start_spec lines 2677 to 2785, preflight and the batch driver; automatic and browser launch paths.
Numbered batch specification JSONserver.py _start_spec lines 2734 to 2740, or _start_resource_pool_spec lines 2657 to 2668Committed batch plan, label, cells and optional Fabric resource informationBatch launcher and driver; report and live helpers in server.py; automatic state stores its relative path.
5. Run Sets/run_tags.jsonserver.py _start_spec lines 2771 to 2780 for automatic launch tags, and tag control handlersTag display, run-set membership and noteTag readers in server.py, tag views in the browser and operators.
5. Run Sets/active_batch.mdBatch driver and control support, read by server active-state helpersRun set, planned cells and status, including stopped_before_completion or aborted_by_operatorserver.py live, _parse_active, control_stop; browser live view.
5. Run Sets/<run set>/execution_log.mdserver.py control_stop lines 2395 to 2404 appends an abort recordTimestamped operator abort sentenceServer reports and operators; the batch closeout path can also read the execution log.
5. Run Sets/<run set>/batch_outcome.jsonBatch driver, read by the server context and tickStatus, reason, planned or attempted cells and completion informationserver.py live and report helpers; auto_experiment.py completion branch at lines 1886 to 1893 and 2103 to 2110.
5. Run Sets/<run set>/closeout_progress.jsonCloseout chain, read by auto_experiment.pyStage keys, status, titles, messages, run id and finish time_discover_closeout_backfill lines 239 to 272; tick_parallel lines 2111 to 2126; server closeout view; app/auto.js.
Unit failure observationsystemd and journal, inspected by server.py _unit_failure_for lines 2506 to 2553Unit, exit_status, result, title, reason, summary and refusallive lines 2145 to 2153; auto_experiment.py through ctx.last_unit_failure; browser live and automatic views.
Active in-memory cachesserver.py catalogue and controller helpersCached catalogue, controller status, live requests and telemetry historyServer endpoints during their cache windows; not a durable campaign record.
Browser local stateapp/index.html shell and browser modulesSelected mode and table layout preferences in local storage; transient DOM statusBrowser modules only. It is not an experiment record.

When save_state writes, it writes a temporary file and replaces the real file at lines 284 to 293. The comment states the atomicity guarantee: readers see either the old complete JSON or the new complete JSON. log_event writes the diary first, then retains the newest 50 entries in history, allowing the History panel to read a bounded copy without scanning the diary.

The record-lineage diagram is a hand-authored structural diagram. It shows the automatic and interface records flowing into run-set records and then into campaign views. It does not claim that the interface writes cell manifests or metric tables.

flowchart LR
    Settings[Browser automatic settings]
    Matrix[cell factor matrix and task library]
    Tick[auto_experiment tick]
    State[auto_experiment_state.json]
    Log[auto_experiment_log.jsonl]
    Draft[draft specification]
    Commit[numbered batch specification]
    Unit[systemd unit and journal]
    Driver[batch driver]
    Run[run-set folder]
    Outcome[batch_outcome.json]
    Closeout[closeout_progress.json]
    Live[/api/live JSON]
    Views[Browser views]
    Fabric[Fabric controller]

    Settings --> State
    Matrix --> Tick
    State --> Tick
    Tick --> Draft
    Tick --> State
    Tick --> Log
    Draft --> Commit
    Commit --> Unit
    Unit --> Driver
    Driver --> Run
    Driver --> Outcome
    Run --> Closeout
    Fabric --> Unit
    State --> Live
    Unit --> Live
    Run --> Live
    Outcome --> Live
    Closeout --> State
    Live --> Views
    Log --> Views

8. The loops and the waits

The server’s automatic loop is unbounded. _auto_tick_loop repeats while True at server.py lines 4472 to 4484. Its wait is 30 seconds, from _AUTO_TICK_SECONDS at line 4437. It stops only when the server task or process is stopped. A tick exception is caught and logged, so the loop continues after that iteration.

The idle watchdog is also unbounded while enabled. _idle_watchdog lines 4520 to 4528 sleeps 60 seconds, checks the time since the last request, and exits the process when the configured idle-minutes threshold has elapsed. It is disabled when --idle-minutes is zero or negative. Browser polling is not owned by the Python loop. The live view requests the endpoint on its front-end interval, documented by the digest as five seconds.

The automatic decision loop has several bounded or time-based waits. Between completed batches it waits for settings["wait_minutes"] at auto_experiment.py lines 1974 to 1986, with the default supplied by settings rather than by the timer. Endpoint failures are quarantined by the automatic state until their recorded retry time. Credit flags expire at the next UTC midnight. Daily launch counts are pruned by _prune_daily_launches during the launch path, retaining the recent accounting window used by the settings rules.

_stop_batch_and_release waits for terminated processes to die for up to 30 seconds according to the server digest and then attempts lease release. _launch_unit waits up to 60 seconds for systemd-run at lines 2446 to 2447. _unit_failure_for waits up to 10 seconds for systemctl and 20 seconds for journalctl at lines 2532 to 2546. The automatic tick itself has no retry loop around its full body. Its retry policy is the next 30 second tick after an exception, with the exception recorded as tick_exception. The endpoint and credit cooldown logic belongs to the state decision and does not retry a refused launch immediately.

tick_parallel has an unbounded endpoint iteration over the current lanes, but each pass is finite: it checks each current unit, records finished results, removes the endpoint, and lets a later planning pass replenish it. The scheduler support in scheduler_lp.py loops over queued jobs and workers until the queue is exhausted or no eligible capacity remains. It uses cooldown timestamps and circuit-breaker state rather than sleeping in the scheduler itself.

9. The guards and refusals

Check and locationRefusal or altered flow
_AUTO_LOOP_ENABLED and startup registration, server.py 4436 to 4453Tests and imports do not start the background loop. A disabled flag prevents task creation.
Path containment and existence, _start_spec 2697 to 2700A path outside the run-set directory or a missing file is rejected with HTTP status 422.
Specification parsing, _start_spec 2701 to 2704An unreadable specification is rejected with HTTP status 422.
Static preflight, _start_spec 2722 to 2732A plan that cannot run as written is refused before systemd launch, unless the caller explicitly skips preflight.
Empty resource-pool cells, _start_resource_pool_spec 2600 to 2603A resource-pool plan with no cells is rejected.
Duplicate resource endpoints, _start_resource_pool_spec 2604 to 2607A resource cannot be selected twice.
Resource model reachability and backend compatibility, _start_resource_pool_spec 2610 to 2627The pool is rejected unless every selected endpoint reports the requested model and backend.
Served identity equality, _start_resource_pool_spec 2628 to 2634A logical plan is rejected when selected endpoints do not serve the same model identity and request policy.
Missing start path, control_start 2789 to 2799The request receives a 422 response and no launch occurs.
Missing resume run set, control_resume 2563 to 2568Resume returns 422.
Missing committed resume specification, control_resume 2569 to 2577Resume returns 422 and does not create a unit.
Current unit active, tick 1873 to 1880Automatic mode records batch_running and does not launch another batch.
Unrelated process active, tick 1952 to 1956Automatic mode records batch_running and waits.
Paused active batch, tick 1958 to 1964Automatic mode records batch_paused and requires resume or abort.
Finishing mode, tick 1966 to 1972The mode becomes off between batches.
Inter-batch wait, tick 1974 to 1986The next launch is held until wait_minutes has elapsed.
Settings validation, tick 1988 to 1992Invalid settings become settings_invalid in hold; no draft is written.
Decision not okay, tick 2001 to 2005No eligible slot or other decision refusal becomes a recorded hold.
Launch result, tick 2059 to 2067Failure becomes launch_failed, is logged and starts the next-attempt wait.
Paid route outcome, tick 1923 to 1939A credit-like failure creates a model credit flag until the next UTC midnight.
Closeout stage status, tick_parallel 2111 to 2126Failed publication creates a closeout backfill warning rather than erasing the lane outcome.
Operator stop target, control_stop 2387 to 2404A missing run-set folder is skipped; an existing target receives an abort line.
Static asset cache, RevalidatingStatic 4511 to 4514Browser assets must revalidate before reuse.
Idle timeout, _idle_watchdog 4520 to 4528The server process exits after the configured period without a request.
Fabric release exception, stop helper 2360 to 2366The response records a failed release string and still returns the rest of the stop report.

The automatic settings validators also enforce the rules in the digest: at least one enabled model, build, task, prompt variant and retry-feedback value; paid-model daily caps in the allowed range; local-model availability on a known endpoint; a request preset for local models; a valid label prefix; at least one endpoint for local models; and an allowed objective and count basis. These are not restated as invented line numbers because their public validator is inside the large automatic module and the digest gives the verified function rather than a narrower source range.

10. The unhappy paths

Each path below has four parts. The first states what the step is supposed to do. The second states why it works that way. The third gives the trigger and the record left. The fourth gives the cost to the driver, scorer, aggregator and interface.

Trigger and four-part accountStatus, exit code and downstream effect
A supplied specification is outside the run-set directory. The start step should launch only an experiment plan owned by this repository. The containment check prevents a request from selecting an arbitrary file. A path that fails the check leaves no batch record and returns error with HTTP 422. The driver never starts, so scorer and aggregator do nothing; the interface shows a refused launch.
The specification is unreadable or fails static preflight. The start step should reject a plan before a background unit hides the reason. The code performs cheap local validation before systemd because the comments identify delayed preflight as an operator-facing failure. An exception or failed gate leaves no run set; it leaves the 422 response with error, refused and possibly gate. The driver, scorer and aggregator receive no work; the interface can show the refusal immediately.
systemd accepts the command and the background unit later dies. The launch step should distinguish command acceptance from a surviving batch. The unit is supervised separately so the server remains responsive. The trigger is a failed or inactive service with a non-success result; _unit_failure_for leaves the unit failure observation with exit_status, result, summary and journal-derived reason. The driver exits, commonly before a run set exists, so no cells reach scoring or aggregation; the live endpoint reports launch_failed and the automatic tick records launch_failed and holds.
The Fabric controller cannot release a lease during pause or stop. The control step should terminate work and release a held resource. The release is attempted explicitly because process termination does not guarantee controller cleanup. A release exception leaves the stop report’s failed-release text and the stopped or aborted records. The driver is stopped, scorer and aggregator see only the partial records already written, and the interface reports the cleanup failure while remaining available.
A unit remains active when the automatic tick runs. The decision step should not create concurrent work for the same serial state. The tick checks systemd state before deciding. The trigger leaves hold with batch_running and preserves current; no new specification or run set is written. The driver continues, scorer and aggregator are unaffected, and the interface shows the active batch.
A batch is stopped before completion. The live step should expose resumability rather than pretending completion. The active marker and process evidence distinguish paused work. The trigger leaves active_batch.md with stopped_before_completion and the partial run-set artefacts. The driver is stopped, scorer and aggregator can see partial data but no complete outcome, and the interface exposes Resume.
The operator presses Stop and abort. The control step should make the pause terminal. The code appends an abort line, writes aborted_by_operator to the active marker, and turns automatic mode off so it cannot relaunch. The trigger leaves execution_log.md, active status and auto_experiment_state.json with mode off. The driver is not resumed, downstream scoring and aggregation stop at available partial records, and the interface removes the resume opportunity.
A batch finishes with a paid-route credit failure. The completion step should prevent an automatic loop from repeatedly attempting a route that could not spend credit. The tick classifies status and failure text, creates a flag with expiry at next UTC midnight and logs it. The trigger leaves credit_flags and a credit_flagged diary event. The driver has already ended, scorer and aggregator retain whatever completed data exists, and the interface shows the model restriction.
A selected endpoint fails or a lane does not create a run set. The parallel completion step should quarantine only the failing endpoint. Endpoint-local accounting makes another endpoint independent. The trigger leaves lane_outcomes, endpoint_failures, a lane_finished event and a retry time. The failed driver contributes no complete lane to scoring or aggregation; the interface shows the lane failure while other lanes can continue.
A publication closeout stage fails. The completion step should retain the experimental outcome while making missing publication work visible. The tick scans closeout_progress.json and records only failed publish_experiment_results stages. The trigger leaves closeout_backfill and a closeout_deferred event. The driver and scorer are already finished, the aggregator data remains, and the interface warns that publication needs backfill.
The state file is missing or damaged. The loop should continue to answer the interface and make a safe next decision. load_state treats missing, invalid and non-dictionary JSON as default_state, then saves a complete replacement after the tick. The trigger leaves a default or repaired auto_experiment_state.json, not a partial parse. No driver is started until valid settings and a decision exist; scorer and aggregator are unaffected; the interface sees off or a hold instead of a server crash.
One tick raises an unexpected exception. The background loop should remain alive for later ticks. The executor boundary catches the exception and appends tick_exception. The trigger leaves the diary event but may leave the prior state unchanged. No new driver is started by that tick, scorer and aggregator are unaffected, and the interface can still answer while the log explains the missed iteration.

The interface itself also has read-side degradation. A missing run-set folder gives the live view no cell rows, a failed Docker or systemd query gives an unavailable or failed observation, and a missing research document is rendered as unavailable rather than invented. Those paths return an API payload or error status and do not create a measured cell record.

11. The metrics produced or fed

This chapter is an interface and orchestration owner. It does not redefine the canonical cell metric columns. It feeds those metrics into views and uses a small number of operational measures for automatic decisions.

Field or measureDestination or displayed columnProgram or pathAdvisory status
Cell factor and planned-cell identityCoverage rows and gap viewsserver.py cells and coverage helpers, lines 100 to 110 and 230 to 270Analytical input, not a replacement for measured metrics.
Cell status and live process stateLive state and cell rowsserver.py live, lines 2117 onwardOperational, not a score.
cells_attemptedAutomatic session iteration and batch outcome displayauto_experiment.py lines 1890 to 1915 and 2106 to 2137Descriptive completion count.
hidden_passesAutomatic iteration historyauto_experiment.py lines 1903 to 1915Descriptive and dependent on the closeout or scoring record.
Batch status and reasonLive, automatic history and lane outcomeServer context, tick lines 1886 to 1893 and tick_parallel lines 2106 to 2130Operational result.
last_decision and hold.reasonAutomatic view and operator historyauto_experiment.py lines 1997 to 2005 and _set_hold lines 1820 to 1824Advisory explanation of why work did or did not launch.
Daily launch count by modelAutomatic settings and cap displayauto_experiment.py lines 2030 to 2034Control accounting, not experiment quality.
Endpoint failure count and retry timeEndpoint availability displaytick_parallel lines 2127 to 2130 and endpoint failure helpersOperational quarantine signal.
Closeout stage statusCloseout warning and backfill viewcloseout_progress.json, read at auto_experiment.py lines 239 to 272 and 2111 to 2126Advisory until publication is repaired.
Unit exit_status and resultLaunch failure panelserver.py _unit_failure_for lines 2532 to 2553Operational failure evidence.
Fabric controller status and catalogueController and model viewsserver catalogue and controller helpers, cached by the serverAvailability and configuration evidence, not a measured outcome.
Timing, token and score fields in cell recordsTiming, evidence, summary and report viewsRead by the server from run-set artefacts; canonical writers are owned by the harness and aggregation chaptersNot advisory when used in the canonical analyses.

The automatic objective can use pooled metrics to compare candidate slots, but the source digest identifies that calculation inside auto_experiment.py, not as a new campaign table. The interface must therefore label launch counts, hold reasons, endpoint failures and publication warnings as operational signals. They explain availability and orchestration; they do not turn an incomplete or failed batch into a valid result.

12. The tests

The verified interface test files in 5. Experiment/10. User Interface/ are test_auto_experiment.py, test_cohorts.py, test_evidence.py, test_evidence_public.py, test_expected_impact.py, test_failure_axes.py, test_landing.py, test_live_cell_states.py, test_live_token_source.py, test_overview_chart.py, test_phd_publish.py, test_proposal_guide.py, test_provider_catalog.py, test_reports.py, test_research_state.py, test_scheduler_lp.py, test_scope_filter.py, test_state_url.py, test_telemetry_throughput.py, test_timeline_phases.py, test_timing_validity_view.py, test_timing_views.py and test_upkeep_view.py. The browser test test_guide_runner.mjs is also present. Their names show coverage of automatic state, cohorts, evidence, live state, publication, reports, research state, scheduler, telemetry and timing views.

The mapping of tests to the exit paths of section 10 was not made. In particular, this review did not establish a test for every distinct 422 refusal, systemd exit status, Fabric release failure, closeout backfill, credit flag, endpoint quarantine or tick exception. The statement is intentional: the repository inventory confirms test files, but it does not provide a verified one-to-one mapping from each unhappy path to a test case.

13. The dated incidents that shaped the code

The comments carry operational history that explains guards which might otherwise look unnecessarily strict.

On 2026-08-31, the interface accepted a batch command and reported success, but the background job refused at preflight and exited two minutes later. The comments in _start_spec at lines 2717 to 2721 explain the addition of the cheap static preflight. The related _unit_failure_for comment at lines 2518 to 2528 explains why the interface reads the specific unit and its journal instead of assuming that accepted means completed. The failure left no run set, so the live view needs an operator-facing explanation.

The same 2026-08-31 comments at lines 2742 to 2753 explain why version-control recording moved into launch_batch.sh. A repository-wide add on a FUSE-mounted disk exceeded the handler’s 300 second limit and left an index lock after the subprocess was killed. The current guard copies the draft into its numbered specification immediately and leaves the slower recording to the background launcher.

The state loader comment records an addendum dated 2026-09-05 at auto_experiment.py lines 216 to 223. Older state files lacked session.iterations. The loader fills that key on every read so a legacy file cannot cause the next append to fail. This is why default_state and load_state are both part of the canonical automatic record path.

The static file comment at server.py lines 4500 to 4507 records a 2026-09-01 browser-cache incident. Changed app.js bytes were served while a browser kept running an older copy. RevalidatingStatic therefore adds no-cache, retaining etag revalidation without forcing a download each time.

The stop comments at server.py lines 2419 to 2422 record the operator ruling that Stop and abort must disable automatic mode as well as terminate the current batch. Otherwise the loop could start another batch a few minutes after the operator had pressed the terminal control.

The parallel comments in auto_experiment.py lines 2072 to 2078 explain endpoint isolation. Each current entry owns a background unit and physical endpoint, so one endpoint can finish or fail without serialising the others. The closeout comments in load_state at lines 239 to 272 explain why unresolved publication failures are replaced by run set rather than duplicated on every read.

The comments in server.py lines 4430 to 4443 explain the shared lock. Without it, a request turning automatic mode off could race a tick that launches a batch and then overwrite the launch record with an older state copy. These dates and reasons are part of the design evidence. They are not a claim that every historical incident has been independently reconstructed beyond the text of the comments.

14. The weakest claim, what was not checked, and the token line

The weakest claim is the completeness of the browser-side reader inventory and the one-to-one test mapping. The interface contains many JavaScript modules and several API response shapes assembled across files. This review verified the owning server functions, the automatic state and log writers, the principal browser entry points and the listed test files, but it did not trace every field from every endpoint into every module. It also did not independently execute the service, systemd unit, Fabric controller, browser suite or publication subprocess. The flow model leaves parallel execution and the exact specification shape partly unmodelled, so those details are reported only where source lines verified them.

Model: openai/gpt-5.6-luna via OpenRouter; tokens: see the job ledger