1. The purpose and the position in the life of the experiment
This mechanism gives an operator a browser view of experiment evidence and gives the experiment an optional autonomous way to choose and start the next batch. A run set is a folder containing the cells and result artefacts for one batch. A cell is one planned programming attempt identified by condition, task, phase, seed and repetition. A batch is the group of cells sent to the batch driver together. The interface is the FastAPI service in 5. Experiment/10. User Interface/server.py; the automatic loop is its background decision service, implemented in auto_experiment.py.
The operator invokes the interface by opening the static application or by sending its HTTP requests. The server reads the campaign matrix, run-set records, monitoring records, cell manifests, controller status and service status. It returns JSON for the browser and writes only the records associated with drafting, tagging, control actions and automatic scheduling. It does not execute a cell itself. It asks a systemd user service, the operating-system supervisor used here for a background process, to run the batch launcher. The driver, scorer and aggregators then create the measured records that the interface displays.
The automatic loop is invoked by the server startup task and, through the auto controls, by the operator. It wakes after thirty seconds, takes a lock shared with requests that change automatic state, and runs one tick. A tick loads persistent state, removes expired failure and credit restrictions, observes any current unit, decides whether work is eligible, writes a draft specification, and sends that specification through the same _start_spec path as a manual launch. Its output is a state file, an append-only event log, and, when a launch is accepted, a systemd unit and a batch specification. Its ordinary purpose is to keep eligible resources working without silently bypassing model, endpoint, wait, credit and daily-count restrictions.
The live page presents five externally visible batch states: idle, starting, running, stopped_resumable, and aborted. stopped_resumable means that a batch was stopped before completion and still has a run set that the operator may resume. aborted is the terminal operator decision that prevents that resume. Automatic mode has four values in its persistent record: off, on, finishing, and hold. finishing means that the current batch may finish but no next batch is to be launched. hold is represented by the hold reason field while the mode remains active and the loop waits for a condition to change.
The position in the life of an experiment is therefore after planning and before execution for a launch, and alongside execution for observation. Manual drafting reaches preflight and the driver through the interface. Automatic drafting reaches the same launch path through a context supplied by the server. During execution the live endpoint reads process, run-set and cell records. After execution the next tick reads the batch outcome and closeout progress, updates the iteration record, and either chooses another batch or records a refusal to proceed. PhD publication is a separate control surface exposed by the same server and is not part of the cell execution path.
2. The reader’s map of the owning files
The principal owner is 5. Experiment/10. User Interface/server.py. Its verified regions are the read endpoints cells at lines 100 to 110 and runsets at lines 113 to 118, coverage and catalogue handling at lines 230 to 350, specification drafting, validation and saving at lines 360 to 510, and live state assembly in live at lines 2117 to 2160 and its continuation through the endpoint body. Control and launch handling is distributed across _stop_batch_and_release at lines 2320 to 2369, control_pause at lines 2372 to 2378, control_stop at lines 2381 to 2437, _launch_unit at lines 2440 to 2458, control_resume at lines 2562 to 2580, _start_resource_pool_spec at lines 2592 to 2674, _start_spec at lines 2677 to 2785, and control_start at lines 2788 to 2799. Automatic settings and state endpoints are at lines 4309 to 4410. The automatic background wiring is _auto_tick_locked at lines 4448 to 4452, _auto_tick_loop at lines 4455 to 4484, static serving at lines 4486 to 4517, the idle watchdog at lines 4520 to 4528, and main at lines 4531 to 4561. The file comments use dated numbered operational changes rather than a single numbered step list. The comments around _start_spec explain static preflight, delayed version-control work and the 2026-08-31 failure it prevents. The comments around _auto_tick_loop explain why an exception is logged without killing the loop.
5. Experiment/10. User Interface/auto_experiment.py is the automatic decision owner. The state region includes default_state at lines 181 to 207, load_state at lines 210 to 236, _discover_closeout_backfill at lines 239 to 272, save_state at lines 284 to 294, append_log at lines 296 to 301, and log_event at lines 303 to 312. The decision and specification helpers occupy the middle of the file, including settings validation, slot construction, coverage selection, decide_next_batch, decide_parallel_batches, and build_spec; their public entry points are verified in the code digest. The numbered comments identify the tick as section 6.2 at lines 1812 to 1818. The main flow is tick at lines 1827 to 2069 and tick_parallel at lines 2072 onward. Finalisation is the branch that fills session.iterations and clears current at lines 1882 to 1950, plus the parallel lane finalisation at lines 2099 to 2140 and its continuation. The file’s own comments call out the state file, the one-action-per-tick rule, session iteration records, endpoint isolation, credit flags and closeout backfill.
5. Experiment/10. User Interface/scheduler_lp.py is the immutable scheduling model used by the interface’s scheduling support. Lines 1 to 37 contain imports and definitions preceding the records. The frozen records include Job, Worker, Assignment, AggregateRecord and PublishRecord in the first 134 lines. The Scheduler flow is plan at lines 135 to 153, complete at lines 155 to 163, aggregate at lines 165 to 173, publish at lines 175 to 182, report_failure at lines 184 to 197, probe at lines 199 to 202, and idle_reason at lines 204 to 233. Helpers are at lines 243 to 287. This file is a scheduling value model, not the writer of the automatic state file.
5. Experiment/10. User Interface/research_state.py supplies the research-state view. Its public functions are verified in the source inventory as load_state, save_state, merge_settings, summary and default_state; the file is the state transformation support used by the interface rather than the automatic launch owner. It reads and normalises research records and returns the summary consumed by the API. No separate persistent record owned by this chapter is asserted for it without a verified writer path.
5. Experiment/10. User Interface/phd_publish.py owns the publication subprocess boundary. Its entry point is main at lines 205 onward. The server exposes publication status and push handlers at lines 3428 to 3433. The browser module app/phd-publish.js calls GET /api/phd-publish/status and POST /api/phd-publish, refreshes every 60 seconds while visible, and delays a post-push refresh by 1.5 seconds. Publication is adjacent to the results interface, but it is not a transition in the automatic batch state machine.
The static shell is 5. Experiment/10. User Interface/app/index.html. It provides the DOM containers, navigation, stylesheets, Plotly and the application modules. The browser modules include app.js, research-state.js, auto.js, timing.js, summary.js, reports.js, evidence.js, cohorts.js, phd-publish.js, and the other view modules listed by the digest. They render server responses and transient browser state. start_ui.sh, stop_ui.sh and publish_ui.sh are operational wrappers. The interface README and the two local specifications explain operation, but this chapter treats source code as the authority for a function, record or line range.
3. The inputs
3.1 Command-line arguments
| Argument | Parser line | Decision made |
|---|---|---|
--host | server.py 4532 to 4534 | Address on which uvicorn binds, default 127.0.0.1. |
--port | server.py 4533 to 4535 | Port on which uvicorn binds, default 8850. |
--idle-minutes | server.py 4535 to 4538 | Whether the idle watchdog is enabled and its timeout. |
The server parser does not define the draft’s claimed --base, --run-sets, --app-dir, --lanes or --auto-loop-enabled arguments, so they are not inputs to this verified design. The browser supplies request JSON instead. control_start reads path and optional skip_preflight at lines 2789 to 2799. control_resume reads run_set and optional skip_preflight at lines 2563 to 2577. Automatic settings arrive through the settings endpoint at lines 4309 onward.
3.2 Environment variables
| Variable | Verified read | What it decides |
|---|---|---|
| Server path and configuration variables | module configuration before the endpoints | The resolved base, run-set, application and launcher paths used by the imported server. The exact variable names are not retained here because a verified source line was not located in the reviewed range. |
| Fabric endpoint and credential variables | _fabric_client and its Fabric client construction, used by _stop_batch_and_release at lines 2360 to 2366 | Controller address and credentials for a lease release. |
| Automatic loop flag | server.py 4436 to 4452 | _AUTO_LOOP_ENABLED, which is a source constant and startup guard, not an environment variable. |
The server source verified for this chapter reads configuration through imported path and Fabric helpers. It does not support the draft’s invented line 3935 environment-variable table. The server’s command line and request bodies are therefore the stable inputs stated here. Secrets are passed to the Fabric client and are not returned in the records described below.
3.3 Files read
| File or family | Verified reader | Decision or view supplied |
|---|---|---|
6. Metrics/cell_factor_matrix.csv | cells and coverage helpers, lines 100 to 110 and 230 to 270 | Planned cells, factors and coverage gaps. |
5. Run Sets/run_index.csv | runsets, lines 113 to 118 | Run-set index shown to the operator. |
7. Monitoring/monitor_state.json | runsets, lines 113 to 118 | Monitoring state. |
5. Run Sets/run_tags.json | tag readers used by _start_spec, lines 2771 to 2780 | Existing tags and automatic tag membership. |
5. Run Sets/cohort_presets.json | cohort API readers | Comparison presets. |
| Task library files | coverage and draft helpers, lines 230 to 270 and 509 onward | Task and chain definitions used for a draft. |
5. Run Sets/<run set>/ artefacts | live lines 2117 to 2160 and report helpers | Cell state, outcomes, closeout and reports. |
5. Run Sets/active_batch.md | _parse_active and live helpers | Current or paused batch marker. |
5. Experiment/10. User Interface/auto_experiment_state.json | auto_experiment.load_state lines 210 to 236 | Automatic mode, settings, current work, restrictions and history. |
5. Experiment/10. User Interface/auto_experiment_log.jsonl | auto_experiment.append_log writes it; the server context exposes the automatic log to the interface | Append-only automatic events. |
| systemd unit state and journal | _unit_failure_for lines 2506 to 2553 | Whether an accepted launch died before producing a run set and its exit status. |
The automatic tick also receives matrix rows, chains, live tasks, process state, active-batch state, unit state, batch outcome and closeout progress through ctx. These are server callables, not direct network inputs to auto_experiment.py.
3.4 Network endpoints
| Endpoint | Client function | Purpose |
|---|---|---|
GET /api/live | browser fetch from the live view | Current batch state, processes, controller and cell information. |
POST /api/control/start | control_start | Validate and launch a supplied specification. |
POST /api/control/pause | control_pause | Stop the current work while retaining resumability where the driver permits it. |
POST /api/control/stop | control_stop | Stop and abort selected run sets and disable automatic mode. |
POST /api/control/resume | control_resume | Launch a resume of a paused run set. |
| Fabric controller API | _fabric_client, called at server.py 2360 to 2366 | Release a lease during operator stop. |
| systemd user manager | _launch_unit, lines 2440 to 2458 | Start the batch driver. |
| systemd user manager and journal | _unit_failure_for, lines 2532 to 2547 | Inspect the accepted unit and retrieve failure text. |
GET /api/phd-publish/status and POST /api/phd-publish | app/phd-publish.js | Inspect and request publication. |
4. The happy path in order
4.1 Server startup
main in 5. Experiment/10. User Interface/server.py lines 4531 to 4561 parses host, port and idle timeout, registers the startup handler, optionally creates the idle watchdog, creates _auto_tick_loop, and starts uvicorn. The code comment gives the design reason for creating the automatic task only here: importing the module or constructing a FastAPI test client must not launch a real batch. The static root at lines 4486 to 4489 serves index.html, and RevalidatingStatic at lines 4492 to 4517 adds Cache-Control: no-cache so a browser checks changed scripts rather than showing an old dashboard.
4.2 Browser assembly and observation
The browser loads app/index.html, its styles and modules. The application modules attach to the DOM and fetch API data. The live module requests /api/live; live begins at line 2117, derives active processes and the run set, and assigns the visible state at lines 2126 to 2139. A process with a run set is running, a process without one is starting, a paused active marker is stopped_resumable, and no such evidence is idle. The response also contains the controller view, generation time and the selected run-set details. The comments give the reason for preferring a process that names its run set over the shared active marker: concurrent endpoint lanes can rewrite that marker independently.
4.3 Draft creation and validation
The operator submits a draft to the specification endpoint at lines 509 onward. The server fills coverage gaps and applies task and chain rules. The save endpoint at lines 667 onward writes a draft specification and returns a launch hint. This is a plan, not an execution record. The operator can then send its path to control_start at lines 2788 to 2799. The handler resolves the path below the base and passes it to _start_spec, keeping HTTP formatting outside the launch body.
4.4 Static preflight and commitment
_start_spec in server.py lines 2677 to 2785 first refuses a path that is not an existing file under the run-set directory, then refuses an unreadable specification. Unless the request skips it, it calls the static preflight at lines 2722 to 2732. The code comment gives the reason: a plan that cannot run should be reported to the person pressing Launch instead of failing minutes later in a background log. A non-numbered draft is copied to a numbered committed specification at lines 2734 to 2740. A tag is attached to the predicted run set at lines 2758 to 2781 after a successful launch request. The prediction is visible if the background unit never creates the folder.
4.5 Systemd launch
_launch_unit in server.py lines 2440 to 2458 constructs a unit name, invokes systemd-run --user with the batch launcher and its arguments, and returns either an error or launched: True with the unit. The subprocess has a 60 second timeout. The code comment distinguishes acceptance of the command from survival of the batch, because a service may pass this boundary and later fail preflight. _start_spec returns the unit, specification and optional tag to either the human endpoint or the automatic context.
4.6 Automatic tick entry
_auto_tick_loop in server.py lines 4455 to 4484 waits for _AUTO_TICK_SECONDS, which is 30 at line 4437, and runs _auto_tick_locked in an executor. _auto_tick_locked lines 4448 to 4452 holds _AUTO_LOCK while it calls auto_experiment.tick through _auto_context. The code comments give both reasons: file and subprocess work must not block the ASGI loop, and a request that changes the same state file must not interleave with a tick. An exception is appended as tick_exception and does not end the loop.
4.7 State recovery and completion handling
tick in auto_experiment.py lines 1827 to 2069 begins by calling load_state at lines 1840 to 1842. load_state lines 210 to 236 treats missing, empty, damaged or non-object JSON as default_state, repairs an old session without iterations, and discovers unresolved publication failures. The tick expires endpoint failures and credit flags, selects serial or parallel operation, and returns immediately with saved state when mode is off at lines 1866 to 1871. If current names a unit, it checks the unit. An active unit causes a batch_running hold at lines 1873 to 1880. A finished unit is matched to its run set and outcome, its open session iteration is closed, a paid-route credit failure may be flagged, and current is cleared at lines 1882 to 1949.
4.8 Eligibility and decision
When no current batch remains, the tick refuses to compete with an unrelated running process at lines 1952 to 1956 and holds when the last batch is paused and resumable at lines 1958 to 1964. finishing becomes off between batches at lines 1966 to 1972. The configured wait interval is enforced at lines 1974 to 1986. validate_settings is called at lines 1988 to 1992. The context then supplies matrix rows, chains and live tasks, and decide_next_batch is called at lines 1994 to 1999. A negative decision becomes a recorded hold at lines 2001 to 2005. This is the point where coverage, enabled models, selected tasks, prompt variants, retry feedback, endpoint availability, caps and previous history affect the next action.
4.9 Specification construction and shared launch
A positive decision is converted by build_spec and written by ctx.write_draft at lines 2007 to 2009. ctx.start_spec invokes the server’s _start_spec, so automatic and manual launches share path checks, static preflight, numbering, tags and systemd. On success, the tick writes current with the unit, label, specification path, predicted run set, launch time, backend, model, request preset, prompt variant, retry feedback and cells at lines 2011 to 2029. It increments daily and session counts, appends an open iteration, clears hold, and logs batch_launched at lines 2030 to 2058. On failure it records launch_failed, sets hold, and sets the last-finished time so repeated failures cannot spin at lines 2059 to 2067. save_state then completes the tick.
4.10 Operator stop and resume
control_pause at lines 2372 to 2378 sends the blocking _stop_batch_and_release work to an executor. That helper terminates units or processes, waits for death, and releases an outstanding Fabric lease at lines 2360 to 2366. control_stop at lines 2381 to 2437 appends an abort line to each targeted run set’s execution_log.md, replaces a paused active marker with aborted_by_operator, and changes automatic mode to off, clearing current work and logging auto_off. control_resume at lines 2562 to 2580 finds the committed specification and starts it with --resume-into. Thus pause preserves a resumable state where the driver permits it, while stop is the explicit terminal action.
5. The state machine
The flow model 5. Experiment/11. Detailed Design/flow-model/flow_model.v001.json does not model this chapter: its open item open-24 lists chapters 01, 02, 07, 11, 12 and 13 as outside the first stage, and it carries no node with a chapter value of 11. The first diagram below is therefore hand-authored from server.py and auto_experiment.py, with the line ranges named on each transition, and its node identifiers (prefixed c11_) are proposed for the next version of the flow model rather than read from it. The proposed identifiers are c11_ui_launch, c11_auto_start, c11_auto_tick, c11_auto_tick_execute, c11_auto_core_tick, c11_auto_decide, c11_auto_build_spec, c11_auto_launch, c11_auto_hold, c11_auto_credit_flag, c11_auto_endpoint_failure, c11_auto_closeout_backfill, c11_auto_save_state, c11_auto_log_event, c11_auto_off and c11_ui_stop. The model supplies the automatic-loop transitions and record witnesses. The live batch states are hand-authored from live and control_stop, because the plan explicitly calls for the interface view and the model seed does not contain those live nodes.
stateDiagram-v2 [*] --> idle idle --> starting: control_start, server.py 2788-2799 idle --> starting: control_resume, server.py 2562-2580 starting --> running: live sees process and run set, server.py 2126-2139 starting --> idle: unit failure, server.py 2145-2153 running --> stopped_resumable: pause completes, server.py 2372-2378 running --> aborted: control_stop, server.py 2381-2437 stopped_resumable --> running: control_resume, server.py 2562-2580 stopped_resumable --> aborted: control_stop, server.py 2387-2417 aborted --> [*]
The live state is an observation derived from process and file evidence, not a separately maintained enum. starting lasts while a process exists without a run-set identity. running requires both. The stopped and aborted values are read from the active marker and the abort record.
The next diagram is also the chapter 11 model subgraph rendered as Mermaid, with the same node identifiers. It names the record-writing nodes that the JSON model places on failure and completion edges. The state labels off, on, finishing and hold are the persistent automatic-loop values. hold is a reason-bearing condition in the state record rather than a distinct mode value in the verified tick code.
stateDiagram-v2 [*] --> off off --> on: c11_auto_start, auto_start endpoint on --> on: c11_auto_tick -> c11_auto_tick_execute on --> hold: c11_auto_decide, no eligible slots hold --> on: later tick finds eligible decision on --> on: c11_auto_build_spec -> c11_auto_launch on --> finishing: c11_auto_off, auto_off endpoint finishing --> off: c11_auto_core_tick, server.py 1943-1949 on --> off: c11_ui_stop, server.py 2423-2437 on --> hold: c11_auto_launch, launch_failed on --> on: c11_auto_credit_flag, credit flag recorded on --> on: c11_auto_endpoint_failure, endpoint failure recorded on --> on: c11_auto_closeout_backfill, closeout warning recorded c11_auto_tick --> c11_auto_tick_execute c11_auto_tick_execute --> c11_auto_core_tick c11_auto_core_tick --> c11_auto_decide c11_auto_decide --> c11_auto_build_spec: decision ok c11_auto_build_spec --> c11_auto_launch c11_auto_launch --> c11_auto_save_state c11_auto_save_state --> c11_auto_log_event c11_auto_log_event --> c11_auto_tick: next 30 second interval
A state transition is not inferred merely from a browser colour. The mode field, current, currents, hold, session.ended_reason, lane_outcomes, credit_flags, endpoint_failures and closeout_backfill are the durable witnesses. A failure may therefore leave the loop active and on hold rather than turning it off.
6. The sequence of one unit of work
A unit of work here is one automatic decision that becomes one supervised batch launch. The diagram is a hand-authored structural diagram from the verified elements in server.py and auto_experiment.py, not a rendering of the flow model. It includes the browser, server event loop, automatic worker thread, state files, systemd and the batch driver. The Fabric controller is shown because the stop path can release its lease, although lease acquisition belongs to the batch driver and its harness chapters.
sequenceDiagram participant Browser as Browser participant Server as server.py participant Loop as asyncio auto loop participant Worker as executor thread participant State as auto_experiment_state.json participant Diary as auto_experiment_log.jsonl participant Systemd as systemd user manager participant Driver as batch driver participant Fabric as Fabric controller Browser->>Server: GET /api/live Server->>State: load automatic state through auto context State-->>Server: mode, current, hold, history Server-->>Browser: live JSON and rendered view Loop->>Loop: sleep 30 seconds Loop->>Worker: run _auto_tick_locked Worker->>State: load_state State-->>Worker: current state or default Worker->>Worker: validate, observe, decide_next_batch Worker->>State: write_draft via server context Worker->>Server: start_spec(draft_path) Server->>Systemd: systemd-run --user batch unit Systemd-->>Server: accepted or nonzero result Server-->>Worker: launched, unit, spec, tag or error alt launch accepted Worker->>State: set current and session iteration Worker->>Diary: append batch_launched Worker->>State: atomic save_state Systemd->>Driver: start batch driver Driver->>Fabric: acquire and use lease as required else launch refused Worker->>State: set hold and last_batch_finished_at Worker->>Diary: append launch_failed Worker->>State: atomic save_state end
The server thread and browser path are separated from the automatic worker by _AUTO_LOCK. The lock covers the state read, decision and write for one tick, but not the browser’s rendering. The tick delegates process and subprocess calls to its context. A unit accepted by systemd is not a completed unit of experiment work. The next tick must observe the service and its run-set outcome before it closes the iteration record.
7. The records
The following table gives the canonical records owned by this chapter. A record is named with its writer and every reader that was verified or is spread through a named reader file. The cell manifests, scores, timing tables and campaign aggregates are written by other chapters. This chapter reads them and therefore does not duplicate their canonical rows.
| Record or table | Writer and line | Fields or contents | Readers |
|---|---|---|---|
5. Experiment/10. User Interface/auto_experiment_state.json | auto_experiment.py, save_state lines 284 to 294 | schema_version, mode, settings, session, current, currents, lane_outcomes, lane_finished_at, endpoint_failures, last_batch_finished_at, credit_flags, closeout_backfill, daily_launches, last_decision, hold, history | auto_experiment.py load_state lines 210 to 236; server.py automatic context and auto endpoints; app/auto.js through the automatic API. |
5. Experiment/10. User Interface/auto_experiment_log.jsonl | auto_experiment.py, append_log lines 296 to 301, called by log_event lines 303 to 312 | One JSON object per line with at, event and detail; the flow model calls these timestamp, event and details | server.py automatic context and API support; app/auto.js History view; operators through the file. |
| Draft specification JSON | server.py automatic context write_draft, and the draft endpoint | Label, cells and selected model, backend, prompt, retry, endpoint and request policy values produced by build_spec | _start_spec lines 2677 to 2785, preflight and the batch driver; automatic and browser launch paths. |
| Numbered batch specification JSON | server.py _start_spec lines 2734 to 2740, or _start_resource_pool_spec lines 2657 to 2668 | Committed batch plan, label, cells and optional Fabric resource information | Batch launcher and driver; report and live helpers in server.py; automatic state stores its relative path. |
5. Run Sets/run_tags.json | server.py _start_spec lines 2771 to 2780 for automatic launch tags, and tag control handlers | Tag display, run-set membership and note | Tag readers in server.py, tag views in the browser and operators. |
5. Run Sets/active_batch.md | Batch driver and control support, read by server active-state helpers | Run set, planned cells and status, including stopped_before_completion or aborted_by_operator | server.py live, _parse_active, control_stop; browser live view. |
5. Run Sets/<run set>/execution_log.md | server.py control_stop lines 2395 to 2404 appends an abort record | Timestamped operator abort sentence | Server reports and operators; the batch closeout path can also read the execution log. |
5. Run Sets/<run set>/batch_outcome.json | Batch driver, read by the server context and tick | Status, reason, planned or attempted cells and completion information | server.py live and report helpers; auto_experiment.py completion branch at lines 1886 to 1893 and 2103 to 2110. |
5. Run Sets/<run set>/closeout_progress.json | Closeout chain, read by auto_experiment.py | Stage keys, status, titles, messages, run id and finish time | _discover_closeout_backfill lines 239 to 272; tick_parallel lines 2111 to 2126; server closeout view; app/auto.js. |
| Unit failure observation | systemd and journal, inspected by server.py _unit_failure_for lines 2506 to 2553 | Unit, exit_status, result, title, reason, summary and refusal | live lines 2145 to 2153; auto_experiment.py through ctx.last_unit_failure; browser live and automatic views. |
| Active in-memory caches | server.py catalogue and controller helpers | Cached catalogue, controller status, live requests and telemetry history | Server endpoints during their cache windows; not a durable campaign record. |
| Browser local state | app/index.html shell and browser modules | Selected mode and table layout preferences in local storage; transient DOM status | Browser modules only. It is not an experiment record. |
When save_state writes, it writes a temporary file and replaces the real file at lines 284 to 293. The comment states the atomicity guarantee: readers see either the old complete JSON or the new complete JSON. log_event writes the diary first, then retains the newest 50 entries in history, allowing the History panel to read a bounded copy without scanning the diary.
The record-lineage diagram is a hand-authored structural diagram. It shows the automatic and interface records flowing into run-set records and then into campaign views. It does not claim that the interface writes cell manifests or metric tables.
flowchart LR Settings[Browser automatic settings] Matrix[cell factor matrix and task library] Tick[auto_experiment tick] State[auto_experiment_state.json] Log[auto_experiment_log.jsonl] Draft[draft specification] Commit[numbered batch specification] Unit[systemd unit and journal] Driver[batch driver] Run[run-set folder] Outcome[batch_outcome.json] Closeout[closeout_progress.json] Live[/api/live JSON] Views[Browser views] Fabric[Fabric controller] Settings --> State Matrix --> Tick State --> Tick Tick --> Draft Tick --> State Tick --> Log Draft --> Commit Commit --> Unit Unit --> Driver Driver --> Run Driver --> Outcome Run --> Closeout Fabric --> Unit State --> Live Unit --> Live Run --> Live Outcome --> Live Closeout --> State Live --> Views Log --> Views
8. The loops and the waits
The server’s automatic loop is unbounded. _auto_tick_loop repeats while True at server.py lines 4472 to 4484. Its wait is 30 seconds, from _AUTO_TICK_SECONDS at line 4437. It stops only when the server task or process is stopped. A tick exception is caught and logged, so the loop continues after that iteration.
The idle watchdog is also unbounded while enabled. _idle_watchdog lines 4520 to 4528 sleeps 60 seconds, checks the time since the last request, and exits the process when the configured idle-minutes threshold has elapsed. It is disabled when --idle-minutes is zero or negative. Browser polling is not owned by the Python loop. The live view requests the endpoint on its front-end interval, documented by the digest as five seconds.
The automatic decision loop has several bounded or time-based waits. Between completed batches it waits for settings["wait_minutes"] at auto_experiment.py lines 1974 to 1986, with the default supplied by settings rather than by the timer. Endpoint failures are quarantined by the automatic state until their recorded retry time. Credit flags expire at the next UTC midnight. Daily launch counts are pruned by _prune_daily_launches during the launch path, retaining the recent accounting window used by the settings rules.
_stop_batch_and_release waits for terminated processes to die for up to 30 seconds according to the server digest and then attempts lease release. _launch_unit waits up to 60 seconds for systemd-run at lines 2446 to 2447. _unit_failure_for waits up to 10 seconds for systemctl and 20 seconds for journalctl at lines 2532 to 2546. The automatic tick itself has no retry loop around its full body. Its retry policy is the next 30 second tick after an exception, with the exception recorded as tick_exception. The endpoint and credit cooldown logic belongs to the state decision and does not retry a refused launch immediately.
tick_parallel has an unbounded endpoint iteration over the current lanes, but each pass is finite: it checks each current unit, records finished results, removes the endpoint, and lets a later planning pass replenish it. The scheduler support in scheduler_lp.py loops over queued jobs and workers until the queue is exhausted or no eligible capacity remains. It uses cooldown timestamps and circuit-breaker state rather than sleeping in the scheduler itself.
9. The guards and refusals
| Check and location | Refusal or altered flow |
|---|---|
_AUTO_LOOP_ENABLED and startup registration, server.py 4436 to 4453 | Tests and imports do not start the background loop. A disabled flag prevents task creation. |
Path containment and existence, _start_spec 2697 to 2700 | A path outside the run-set directory or a missing file is rejected with HTTP status 422. |
Specification parsing, _start_spec 2701 to 2704 | An unreadable specification is rejected with HTTP status 422. |
Static preflight, _start_spec 2722 to 2732 | A plan that cannot run as written is refused before systemd launch, unless the caller explicitly skips preflight. |
Empty resource-pool cells, _start_resource_pool_spec 2600 to 2603 | A resource-pool plan with no cells is rejected. |
Duplicate resource endpoints, _start_resource_pool_spec 2604 to 2607 | A resource cannot be selected twice. |
Resource model reachability and backend compatibility, _start_resource_pool_spec 2610 to 2627 | The pool is rejected unless every selected endpoint reports the requested model and backend. |
Served identity equality, _start_resource_pool_spec 2628 to 2634 | A logical plan is rejected when selected endpoints do not serve the same model identity and request policy. |
Missing start path, control_start 2789 to 2799 | The request receives a 422 response and no launch occurs. |
Missing resume run set, control_resume 2563 to 2568 | Resume returns 422. |
Missing committed resume specification, control_resume 2569 to 2577 | Resume returns 422 and does not create a unit. |
Current unit active, tick 1873 to 1880 | Automatic mode records batch_running and does not launch another batch. |
Unrelated process active, tick 1952 to 1956 | Automatic mode records batch_running and waits. |
Paused active batch, tick 1958 to 1964 | Automatic mode records batch_paused and requires resume or abort. |
Finishing mode, tick 1966 to 1972 | The mode becomes off between batches. |
Inter-batch wait, tick 1974 to 1986 | The next launch is held until wait_minutes has elapsed. |
Settings validation, tick 1988 to 1992 | Invalid settings become settings_invalid in hold; no draft is written. |
Decision not okay, tick 2001 to 2005 | No eligible slot or other decision refusal becomes a recorded hold. |
Launch result, tick 2059 to 2067 | Failure becomes launch_failed, is logged and starts the next-attempt wait. |
Paid route outcome, tick 1923 to 1939 | A credit-like failure creates a model credit flag until the next UTC midnight. |
Closeout stage status, tick_parallel 2111 to 2126 | Failed publication creates a closeout backfill warning rather than erasing the lane outcome. |
Operator stop target, control_stop 2387 to 2404 | A missing run-set folder is skipped; an existing target receives an abort line. |
Static asset cache, RevalidatingStatic 4511 to 4514 | Browser assets must revalidate before reuse. |
Idle timeout, _idle_watchdog 4520 to 4528 | The server process exits after the configured period without a request. |
| Fabric release exception, stop helper 2360 to 2366 | The response records a failed release string and still returns the rest of the stop report. |
The automatic settings validators also enforce the rules in the digest: at least one enabled model, build, task, prompt variant and retry-feedback value; paid-model daily caps in the allowed range; local-model availability on a known endpoint; a request preset for local models; a valid label prefix; at least one endpoint for local models; and an allowed objective and count basis. These are not restated as invented line numbers because their public validator is inside the large automatic module and the digest gives the verified function rather than a narrower source range.
10. The unhappy paths
Each path below has four parts. The first states what the step is supposed to do. The second states why it works that way. The third gives the trigger and the record left. The fourth gives the cost to the driver, scorer, aggregator and interface.
| Trigger and four-part account | Status, exit code and downstream effect |
|---|---|
A supplied specification is outside the run-set directory. The start step should launch only an experiment plan owned by this repository. The containment check prevents a request from selecting an arbitrary file. A path that fails the check leaves no batch record and returns error with HTTP 422. The driver never starts, so scorer and aggregator do nothing; the interface shows a refused launch. | |
The specification is unreadable or fails static preflight. The start step should reject a plan before a background unit hides the reason. The code performs cheap local validation before systemd because the comments identify delayed preflight as an operator-facing failure. An exception or failed gate leaves no run set; it leaves the 422 response with error, refused and possibly gate. The driver, scorer and aggregator receive no work; the interface can show the refusal immediately. | |
systemd accepts the command and the background unit later dies. The launch step should distinguish command acceptance from a surviving batch. The unit is supervised separately so the server remains responsive. The trigger is a failed or inactive service with a non-success result; _unit_failure_for leaves the unit failure observation with exit_status, result, summary and journal-derived reason. The driver exits, commonly before a run set exists, so no cells reach scoring or aggregation; the live endpoint reports launch_failed and the automatic tick records launch_failed and holds. | |
| The Fabric controller cannot release a lease during pause or stop. The control step should terminate work and release a held resource. The release is attempted explicitly because process termination does not guarantee controller cleanup. A release exception leaves the stop report’s failed-release text and the stopped or aborted records. The driver is stopped, scorer and aggregator see only the partial records already written, and the interface reports the cleanup failure while remaining available. | |
A unit remains active when the automatic tick runs. The decision step should not create concurrent work for the same serial state. The tick checks systemd state before deciding. The trigger leaves hold with batch_running and preserves current; no new specification or run set is written. The driver continues, scorer and aggregator are unaffected, and the interface shows the active batch. | |
A batch is stopped before completion. The live step should expose resumability rather than pretending completion. The active marker and process evidence distinguish paused work. The trigger leaves active_batch.md with stopped_before_completion and the partial run-set artefacts. The driver is stopped, scorer and aggregator can see partial data but no complete outcome, and the interface exposes Resume. | |
The operator presses Stop and abort. The control step should make the pause terminal. The code appends an abort line, writes aborted_by_operator to the active marker, and turns automatic mode off so it cannot relaunch. The trigger leaves execution_log.md, active status and auto_experiment_state.json with mode off. The driver is not resumed, downstream scoring and aggregation stop at available partial records, and the interface removes the resume opportunity. | |
A batch finishes with a paid-route credit failure. The completion step should prevent an automatic loop from repeatedly attempting a route that could not spend credit. The tick classifies status and failure text, creates a flag with expiry at next UTC midnight and logs it. The trigger leaves credit_flags and a credit_flagged diary event. The driver has already ended, scorer and aggregator retain whatever completed data exists, and the interface shows the model restriction. | |
A selected endpoint fails or a lane does not create a run set. The parallel completion step should quarantine only the failing endpoint. Endpoint-local accounting makes another endpoint independent. The trigger leaves lane_outcomes, endpoint_failures, a lane_finished event and a retry time. The failed driver contributes no complete lane to scoring or aggregation; the interface shows the lane failure while other lanes can continue. | |
A publication closeout stage fails. The completion step should retain the experimental outcome while making missing publication work visible. The tick scans closeout_progress.json and records only failed publish_experiment_results stages. The trigger leaves closeout_backfill and a closeout_deferred event. The driver and scorer are already finished, the aggregator data remains, and the interface warns that publication needs backfill. | |
The state file is missing or damaged. The loop should continue to answer the interface and make a safe next decision. load_state treats missing, invalid and non-dictionary JSON as default_state, then saves a complete replacement after the tick. The trigger leaves a default or repaired auto_experiment_state.json, not a partial parse. No driver is started until valid settings and a decision exist; scorer and aggregator are unaffected; the interface sees off or a hold instead of a server crash. | |
One tick raises an unexpected exception. The background loop should remain alive for later ticks. The executor boundary catches the exception and appends tick_exception. The trigger leaves the diary event but may leave the prior state unchanged. No new driver is started by that tick, scorer and aggregator are unaffected, and the interface can still answer while the log explains the missed iteration. |
The interface itself also has read-side degradation. A missing run-set folder gives the live view no cell rows, a failed Docker or systemd query gives an unavailable or failed observation, and a missing research document is rendered as unavailable rather than invented. Those paths return an API payload or error status and do not create a measured cell record.
11. The metrics produced or fed
This chapter is an interface and orchestration owner. It does not redefine the canonical cell metric columns. It feeds those metrics into views and uses a small number of operational measures for automatic decisions.
| Field or measure | Destination or displayed column | Program or path | Advisory status |
|---|---|---|---|
| Cell factor and planned-cell identity | Coverage rows and gap views | server.py cells and coverage helpers, lines 100 to 110 and 230 to 270 | Analytical input, not a replacement for measured metrics. |
| Cell status and live process state | Live state and cell rows | server.py live, lines 2117 onward | Operational, not a score. |
cells_attempted | Automatic session iteration and batch outcome display | auto_experiment.py lines 1890 to 1915 and 2106 to 2137 | Descriptive completion count. |
hidden_passes | Automatic iteration history | auto_experiment.py lines 1903 to 1915 | Descriptive and dependent on the closeout or scoring record. |
Batch status and reason | Live, automatic history and lane outcome | Server context, tick lines 1886 to 1893 and tick_parallel lines 2106 to 2130 | Operational result. |
last_decision and hold.reason | Automatic view and operator history | auto_experiment.py lines 1997 to 2005 and _set_hold lines 1820 to 1824 | Advisory explanation of why work did or did not launch. |
| Daily launch count by model | Automatic settings and cap display | auto_experiment.py lines 2030 to 2034 | Control accounting, not experiment quality. |
| Endpoint failure count and retry time | Endpoint availability display | tick_parallel lines 2127 to 2130 and endpoint failure helpers | Operational quarantine signal. |
| Closeout stage status | Closeout warning and backfill view | closeout_progress.json, read at auto_experiment.py lines 239 to 272 and 2111 to 2126 | Advisory until publication is repaired. |
Unit exit_status and result | Launch failure panel | server.py _unit_failure_for lines 2532 to 2553 | Operational failure evidence. |
| Fabric controller status and catalogue | Controller and model views | server catalogue and controller helpers, cached by the server | Availability and configuration evidence, not a measured outcome. |
| Timing, token and score fields in cell records | Timing, evidence, summary and report views | Read by the server from run-set artefacts; canonical writers are owned by the harness and aggregation chapters | Not advisory when used in the canonical analyses. |
The automatic objective can use pooled metrics to compare candidate slots, but the source digest identifies that calculation inside auto_experiment.py, not as a new campaign table. The interface must therefore label launch counts, hold reasons, endpoint failures and publication warnings as operational signals. They explain availability and orchestration; they do not turn an incomplete or failed batch into a valid result.
12. The tests
The verified interface test files in 5. Experiment/10. User Interface/ are test_auto_experiment.py, test_cohorts.py, test_evidence.py, test_evidence_public.py, test_expected_impact.py, test_failure_axes.py, test_landing.py, test_live_cell_states.py, test_live_token_source.py, test_overview_chart.py, test_phd_publish.py, test_proposal_guide.py, test_provider_catalog.py, test_reports.py, test_research_state.py, test_scheduler_lp.py, test_scope_filter.py, test_state_url.py, test_telemetry_throughput.py, test_timeline_phases.py, test_timing_validity_view.py, test_timing_views.py and test_upkeep_view.py. The browser test test_guide_runner.mjs is also present. Their names show coverage of automatic state, cohorts, evidence, live state, publication, reports, research state, scheduler, telemetry and timing views.
The mapping of tests to the exit paths of section 10 was not made. In particular, this review did not establish a test for every distinct 422 refusal, systemd exit status, Fabric release failure, closeout backfill, credit flag, endpoint quarantine or tick exception. The statement is intentional: the repository inventory confirms test files, but it does not provide a verified one-to-one mapping from each unhappy path to a test case.
13. The dated incidents that shaped the code
The comments carry operational history that explains guards which might otherwise look unnecessarily strict.
On 2026-08-31, the interface accepted a batch command and reported success, but the background job refused at preflight and exited two minutes later. The comments in _start_spec at lines 2717 to 2721 explain the addition of the cheap static preflight. The related _unit_failure_for comment at lines 2518 to 2528 explains why the interface reads the specific unit and its journal instead of assuming that accepted means completed. The failure left no run set, so the live view needs an operator-facing explanation.
The same 2026-08-31 comments at lines 2742 to 2753 explain why version-control recording moved into launch_batch.sh. A repository-wide add on a FUSE-mounted disk exceeded the handler’s 300 second limit and left an index lock after the subprocess was killed. The current guard copies the draft into its numbered specification immediately and leaves the slower recording to the background launcher.
The state loader comment records an addendum dated 2026-09-05 at auto_experiment.py lines 216 to 223. Older state files lacked session.iterations. The loader fills that key on every read so a legacy file cannot cause the next append to fail. This is why default_state and load_state are both part of the canonical automatic record path.
The static file comment at server.py lines 4500 to 4507 records a 2026-09-01 browser-cache incident. Changed app.js bytes were served while a browser kept running an older copy. RevalidatingStatic therefore adds no-cache, retaining etag revalidation without forcing a download each time.
The stop comments at server.py lines 2419 to 2422 record the operator ruling that Stop and abort must disable automatic mode as well as terminate the current batch. Otherwise the loop could start another batch a few minutes after the operator had pressed the terminal control.
The parallel comments in auto_experiment.py lines 2072 to 2078 explain endpoint isolation. Each current entry owns a background unit and physical endpoint, so one endpoint can finish or fail without serialising the others. The closeout comments in load_state at lines 239 to 272 explain why unresolved publication failures are replaced by run set rather than duplicated on every read.
The comments in server.py lines 4430 to 4443 explain the shared lock. Without it, a request turning automatic mode off could race a tick that launches a batch and then overwrite the launch record with an older state copy. These dates and reasons are part of the design evidence. They are not a claim that every historical incident has been independently reconstructed beyond the text of the comments.
14. The weakest claim, what was not checked, and the token line
The weakest claim is the completeness of the browser-side reader inventory and the one-to-one test mapping. The interface contains many JavaScript modules and several API response shapes assembled across files. This review verified the owning server functions, the automatic state and log writers, the principal browser entry points and the listed test files, but it did not trace every field from every endpoint into every module. It also did not independently execute the service, systemd unit, Fabric controller, browser suite or publication subprocess. The flow model leaves parallel execution and the exact specification shape partly unmodelled, so those details are reported only where source lines verified them.
Model: openai/gpt-5.6-luna via OpenRouter; tokens: see the job ledger