1. The purpose and the position in the life of the experiment
An agent back end is the harness component that gives a coding agent a measured place to run, and a prompt is the instruction text placed before the task that the agent must carry out. Together they carry one maintenance attempt from the cell runner to the agent and bring back the attempt evidence. The cell runner, run_cell.py, invokes this mechanism once for each pass. It produces the final prompt argument, the agent process result, the transcript and error files, the handoff evidence, and, for a Fabric route, the lease lifecycle that explains how the model service was used.
A Fabric lease is a temporary permission from the model controller to send requests for one pass. An admission is the controller request made before that permission exists. A heartbeat is a periodic renewal signal while the permission is in use. The two Fabric back ends, Codex and Claude, share this lifecycle but use different agent command lines and protocols. The ordinary Claude route can run without Fabric and remains the non-Fabric alternative. The back end does not score the product change. It starts and stops the agent, preserves its outputs, reports process status and usage where its transcript provides it, and leaves the runner to validate the handoff and score the cell.
The prompt side is invoked inside the pass loop. run_cell.py combines the one prompt of record, the task package, and pass overhead. The task package is the task brief plus visible verification instructions. The overhead contains the pass header and, when applicable, a retry notice. A retry notice is the short instruction shown after a preceding failure. Its visible form quotes the visible failure summary, its environment form explains a location-dependent failure, and its hidden forms point back to the task requirements without naming hidden files or checks. The resulting string is passed as one command argument and the three source portions are recorded as injection categories by character count.
After the process returns, the runner captures HANDOFF.json before checking anything else and sends the structured document through the handoff library. A handoff is the agent’s JSON closing note about its work. handoff_lib.py loads it, validates its required shape against the schema or a manual equivalent, and returns problems as data. A malformed or missing handoff does not by itself change the cell verdict. The flow-model subgraph for this chapter names the prompt-received node agent.prompt_received, the request loop agent.request, the reply agent.reply, the tool actions agent.tool_call.read, agent.tool_call.search, agent.tool_call.edit, agent.tool_call.test_run and agent.tool_call.other, the result agent.tool_result, compaction agent.compaction, notice agent.notice, handoff agent.handoff_written, pass end agent.pass_ended, and the retry nodes agent.saw_visible_failure and agent.saw_hidden_feedback.
2. The reader’s map of the owning files
The owning scripts are 5. Experiment/1. Harness/scripts/agent_backends.py, 5. Experiment/1. Harness/scripts/handoff_lib.py, and 5. Experiment/1. Harness/scripts/run_cell.py. The first is over 50,000 bytes, so its own numbered and named implementation regions are used below rather than treating it as a short module.
2.1 agent_backends.py
The constants and small pure builders occupy lines 22 to 123. They define backend names, model defaults, interface requirements, context defaults and the lease retry setting. Compatibility and model-window helpers occupy lines 133 to 358, including fabric_backend_compatibility, endpoint qualification, catalogue caching, compaction and output-cap derivation. The result record and backend selector are at lines 361 to 379. Command builders and container cleanup occupy lines 382 to 542. Transcript parsers occupy lines 545 to 623. Event filtering occupies lines 646 to 661.
The shared main flow is the _FabricLeaseBackend class at lines 664 to 1061. Its constructor and model qualification are at lines 677 to 738. Its subclass hooks are at lines 748 to 765. Lease acquisition and its retry loop are at lines 768 to 827. run_pass starts at line 829 and covers lifecycle construction, event recording, qualification, preparation, admission, environment application, command execution, interruption handling and cleanup through line 1061. The concrete Codex back end is at lines 1064 to 1098. The concrete Claude back end is at lines 1101 to 1160. Finalisation is therefore split between the shared finally block at lines 1005 to 1047 and result construction at lines 1049 to 1061, with backend-specific error and usage extraction at lines 1092 to 1098 and 1147 to 1160.
2.2 handoff_lib.py
The constants and file loader are at lines 10 to 32. load reads JSON and returns either a dictionary or None. The validation helpers are at lines 35 to 82. Manual validation of the handoff shape is at lines 84 to 157. The public validate entry point is at lines 160 to 181 and prefers jsonschema when installed, falling back to the manual validator when it is not. The compact result formatter summary_line is at lines 183 to 190. There is no network path in this file and no finalisation region beyond returning validation text.
2.3 run_cell.py
The prompt constants and template names are at lines 81 to 140. The retry constants and frozen-feedback loader are at lines 178 to 285, with unmet_criterion at lines 291 to 313. The prompt and retry composition helpers are at lines 869 to 943, specifically compose_artifact_hint at lines 869 to 922 and compose_retry_notice at lines 925 to 943. The task package builder is at lines 1560 to 1581. The argument parser defines the relevant command-line inputs at lines 1516 to 1553. Runtime resolution of the Fabric route and its environment guards is at lines 1681 to 1755. The per-pass prompt and backend invocation are at lines 2464 to 2547. Handoff capture and subsequent result accounting begin at lines 2549 to 2555 and continue in the runner’s finalisation path.
3. The inputs
The command-line arguments table names the parser location, not a guessed caller interface.
| Argument | Parser line | Decision made |
|---|---|---|
--model | run_cell.py 1516 | Selects the model alias, otherwise the Fabric model environment or backend default is used. |
--prompt-variant | run_cell.py 1517 to 1520 | Selects none or artifact_hint. |
--retry-feedback | run_cell.py 1523 to 1525 | Enables or disables opening the frozen retry map. |
--agent-backend | run_cell.py 1527 to 1530 | Selects the ordinary Claude, DGX Codex or DGX Claude route. |
--fabric-endpoint | run_cell.py 1543 to 1546 | Selects the named Fabric endpoint when a Fabric backend is active. |
--start-tree | run_cell.py 1553 to 1555 | Supplies a chained starting tree to the runner’s surrounding pass logic. |
The environment table includes variables directly read by these owning regions and variables set around the child process.
| Variable | Read or applied at | Decision made |
|---|---|---|
FIVEB_LEASE_RETRY_SECONDS | agent_backends.py 113 to 123, consumed at 795 | Positive wait between retryable lease failures, default 60 seconds. |
DGX_SPARK_FABRIC_MODEL | run_cell.py 1682 to 1683 | Fabric model alias when --model is absent. |
FIVEB_CELL_CONTAINER | run_cell.py 1702 to 1717 | Required image for either DGX backend. |
FIVEB_FABRIC_ENDPOINT | run_cell.py 1719 to 1721 | Named endpoint fallback when the argument is absent. |
DGX_SPARK_FABRIC_LANE | run_cell.py 1724 to 1732 | Required and restricted admission lane. |
ANTHROPIC_API_KEY, FIVEB_CELL_CREDENTIALS | run_cell.py 1734 to 1749 | Their presence refuses the measured DGX Claude route. |
DGX_SPARK_FABRIC_TOKEN | agent_backends.py 420 to 422 | Forwarded by name to Codex container execution. |
DGX_SPARK_FABRIC_LEASE_ID | agent_backends.py 420 to 422 and 755 to 757 | Forwarded by name and bound to the active lease. |
ANTHROPIC_AUTH_TOKEN, ANTHROPIC_CUSTOM_HEADERS | agent_backends.py 1140 to 1145 and 497 to 500 | Forwarded by name for the Claude Fabric request and lease header. |
FIVEB_CHAINED_GATE_SATISFIED | run_cell.py 946 to 950 | Restores the earlier chained-step refusal when set to 0. |
The files-read table distinguishes prompt inputs from execution evidence.
| File | Read at | Decision made |
|---|---|---|
prompts/maintenance.single_agent.v006.md | run_cell.py prompt-loading path beginning at 81 and used by the pass loop at 2480 | Supplies the prompt of record. |
prompts/artifact_hint.v001.md | run_cell.py 907 to 919 | Supplies the optional hint body after the staged inventory passes its filters. |
lap_artifact_manifest.json in the workspace | run_cell.py 881 to 901 | Supplies candidate documentation paths and inline markers for the hint. |
task.md or fallback TICKET.md | run_cell.py 1570 to 1578 | Supplies the task brief. |
visible_tests.md or fallback HOW_TO_VERIFY.md | run_cell.py 1573 to 1580 | Supplies visible verification instructions. |
retry_feedback.json | run_cell.py 271 to 285 and 2446 to 2447 | Maps a hidden check to a task criterion when retry feedback is enabled. |
| Codex profile template | agent_backends.py 396 to 405 and DgxCodexBackend._prepare 1078 to 1082 | Supplies non-secret provider configuration before the Codex container starts. |
| Agent transcript JSON lines | agent_backends.py 553 to 605 and 611 to 623 | Supplies Codex usage and process-level error details. |
| Handoff JSON | handoff_lib.py 26 to 32 and run_cell.py handoff capture beginning at 2549 | Supplies the structured closing note for validation. |
The network table covers the external calls made by the back ends. The prompt composer and handoff validator make no network calls.
| Endpoint or operation | Client function | Decision made |
|---|---|---|
| Fabric catalogue | FabricClient.catalog, called from agent_backends.py 726 to 738 | Qualifies the selected alias and obtains a context window on a non-DGX endpoint. |
| Fabric admission and polling | FabricClient.acquire, called through _acquire_lease at agent_backends.py 799 to 800 | Obtains or waits for the lease. |
| Fabric heartbeat | HeartbeatTask, started at agent_backends.py 977 to 978 | Keeps the lease alive while the child runs. |
| Fabric release | FabricClient.release, agent_backends.py 1027 to 1033 | Returns a claimed lease with the outcome and detail. |
| Fabric cancellation | FabricClient.cancel_admission, agent_backends.py 1037 to 1045 | Removes scheduler state when admission exists without a lease. |
| Agent model route | The Codex or Claude command inside the container, composed at agent_backends.py 409 to 513 | Sends the composed prompt through the selected provider route. |
4. The happy path in order
4.1 Backend selection
The runner first selects a route with select_backend in agent_backends.py lines 373 to 379. It strips the supplied name, defaults to the ordinary Claude route, and refuses any name not in the three backend constants. This is deliberately a fail-fast boundary: later code can construct a command only for a known route.
4.2 Prompt source loading
The runner obtains the single prompt of record and the task package before the pass body. compose_task_package in run_cell.py lines 1560 to 1581 reads the task brief and visible verification file, preferring the additive names only when the retired names are absent. If the selected variant is artifact hint, compose_artifact_hint in run_cell.py lines 869 to 922 reads the workspace inventory, rejects test-tree paths and the inventory itself, keeps existing documentation suffixes, and renders the remaining paths. The reason in the code comments is experimental control: the variant adds a recorded, switchable paragraph rather than replacing the prompt or silently handing over a reading list.
4.3 Pass prompt composition
At the start of each pass, the loop in run_cell.py lines 2464 to 2480 writes a pass header and selects a retry notice when the pass is not the first or a prior failure exists. compose_retry_notice in lines 925 to 943 chooses the visible, environment or hidden form. A hidden failure can be resolved to a criterion by unmet_criterion in lines 291 to 313, but absent or malformed feedback falls back to the generic notice. The prompt template, task package and overhead are joined at line 2480. The independent length guard at lines 2481 to 2489 bounds the overhead when the single argument would exceed MAX_PROMPT_CHARS; it does not invent a replacement task. The injection record at lines 2490 to 2496 preserves the three source categories and their character counts.
4.4 Model and endpoint qualification
For a Fabric route, run_pass in agent_backends.py lines 829 to 1061 creates the lifecycle record and client, then calls _qualify_selected_endpoint_model at lines 914 to 918. The named endpoint path reads its catalogue, finds the exact alias and checks its interface and capabilities at lines 726 to 738. The original dgx path uses the reviewed context-window lookup at lines 723 to 725. The reason for placing the check inside the pass boundary is stated at lines 914 to 917: a refusal must leave an unstarted pass record rather than escape without an error artifact.
4.5 Admission and lease acquisition
The backend calls _acquire_lease in agent_backends.py lines 944 to 957, supplying the lane, workload, alias, priority, lease length, stable idempotency key, structured cell metadata, requirements and any saved lease. _acquire_lease begins at line 768 and loops from lines 795 to 827. It returns the controller result immediately when admission succeeds. A retryable exception is recorded as a waiting event, followed by the configured sleep. The reason in the docstring at lines 774 to 786 is that maintenance, restart and temporary capacity failures may clear, while invalid credentials or an unknown alias cannot.
4.6 Container preparation and command construction
The Codex subclass calls _prepare at agent_backends.py lines 1078 to 1082, rendering a cell-local TOML profile. The shared pass applies _pass_env at lines 971 to 974, then calls _compose_argv at lines 975 to 976. Codex uses codex_container_argv at lines 409 to 434. Claude uses DgxClaudeBackend._compose_argv at lines 1120 to 1138, removes an earlier pass CID file, remembers the new path and calls claude_fabric_container_argv at lines 437 to 513. A CID file is a host file containing the Docker container identifier. The comments at lines 448 to 460 explain why credentials are passed by environment-variable name and why telemetry and unrelated feature traffic are disabled. Context, output and compaction values are supplied only when their provenance is available at lines 477 to 489.
4.7 Agent execution and heartbeat
run_pass starts a HeartbeatTask at agent_backends.py lines 977 to 978 and calls the injected runner at lines 979 to 983. The runner writes the raw transcript and stderr path. The heartbeat gives the controller a live lease signal while the agent process runs. The process status becomes completed only when the runner returns zero at lines 988 to 989. A nonzero status remains failed, although cleanup still attempts to stop the container and release or cancel controller state.
4.8 Finalisation and handoff validation
The shared cleanup at agent_backends.py lines 1005 to 1047 removes every temporary environment variable, stops an abnormal container through its CID file, stops the heartbeat, releases a lease or cancels an unclaimed admission, and records the final lifecycle. AgentRunResult is built at lines 1049 to 1061. Codex usage is parsed by parse_codex_usage in lines 545 to 605; Claude returns no separate provider usage because its registered ledger parser owns that attribution at lines 1157 to 1160. Back in run_cell.py, capture_handoff begins at lines 1156 to 1193 and is called after the backend returns at lines 2549 to 2555. load and validate in handoff_lib.py lines 26 to 32 and 160 to 181 make the handoff safe to consume without allowing validation failure to become a new cell verdict.
5. The state machine
The following diagram is the chapter 05 subgraph of 5. Experiment/11. Detailed Design/flow-model/flow_model.v001.json, rendered as Mermaid. The node identifiers are agent.prompt_received, agent.request, agent.reply, agent.tool_call.read, agent.tool_call.search, agent.tool_call.edit, agent.tool_call.test_run, agent.tool_call.other, agent.tool_result, agent.compaction, agent.notice, agent.handoff_written, agent.pass_ended, agent.saw_visible_failure and agent.saw_hidden_feedback. The model describes the agent-visible pass states and their evidence transitions. The backend’s controller lifecycle is represented in the same pass by the lifecycle values admission_state, lease_id, claimed_at, last_heartbeat_at, released_at and interrupted_at, rather than by unverified synthetic state names.
stateDiagram-v2 [*] --> agent_prompt_received: run_cell.py 2472 to 2480 agent_prompt_received --> agent_request: prompt passed to agent, run_cell.py 2498 to 2520 agent_request --> agent_reply: model response, ledger_lib.py 420 to 539 agent_reply --> agent_tool_call_read: classified read, ledger_lib.py 169 to 200 agent_reply --> agent_tool_call_search: classified search, ledger_lib.py 169 to 200 agent_reply --> agent_tool_call_edit: classified edit, ledger_lib.py 169 to 200 agent_reply --> agent_tool_call_test_run: classified test, ledger_lib.py 169 to 200 agent_reply --> agent_tool_call_other: other tool call, ledger_lib.py 169 to 200 agent_tool_call_read --> agent_tool_result: timing_lib.py 818 to 830 agent_tool_call_search --> agent_tool_result: timing_lib.py 818 to 830 agent_tool_call_edit --> agent_tool_result: timing_lib.py 818 to 830 agent_tool_call_test_run --> agent_tool_result: timing_lib.py 818 to 830 agent_tool_call_other --> agent_tool_result: timing_lib.py 818 to 830 agent_tool_result --> agent_request: next request, bound by agent tool agent_reply --> agent_compaction: compact_boundary, timing_lib.py 445 agent_reply --> agent_notice: system line, timing_lib.py 445 agent_request --> agent_handoff_written: closing note, run_cell.py 1156 to 1193 agent_handoff_written --> agent_pass_ended: result line, run_cell.py 1463 to 1500 agent_pass_ended --> agent_saw_visible_failure: next pass notice, run_cell.py 925 to 943 agent_pass_ended --> agent_saw_hidden_feedback: next pass notice, run_cell.py 925 to 943
The backend record has a simpler lifecycle. Before admission it has no lease identifier. Admission adds an admission identifier and state. Claim or recovery adds lease_id and model provenance. Heartbeat updates last_heartbeat_at. Normal release adds released_at; interruption adds interrupted_at; an admission without a lease is cancelled. These are record values observed in run_pass lines 835 to 899 and cleanup lines 1005 to 1047, so they are not presented as a separate controller taxonomy.
6. The sequence of one unit of work
One unit is one pass of one cell. The following diagram names programs, the runner thread, the external Fabric service, the Docker child, and the files. It also shows the prompt and ledger path. It is the chapter 05 flow-model sequence rendered as Mermaid, with the named chapter nodes agent.prompt_received, agent.request, agent.reply, agent.tool_result, agent.handoff_written and agent.pass_ended placed at their verified code boundaries.
sequenceDiagram participant Runner as run_cell.py runner participant Backend as agent_backends.py participant Fabric as Fabric controller participant Docker as agent container participant Transcript as pass stream and stderr files participant Ledger as turn and token ledger Runner->>Runner: compose_task_package, run_cell.py 1560 to 1581 Runner->>Runner: agent.prompt_received, run_cell.py 2472 to 2496 Runner->>Backend: run_pass, agent_backends.py 2536 to 2541 Backend->>Fabric: catalog and admission, agent_backends.py 726 to 800 Fabric-->>Backend: claim or retryable refusal Backend->>Docker: composed Docker argv, agent_backends.py 975 to 978 Backend->>Fabric: heartbeat, HeartbeatTask Docker->>Fabric: provider requests through lease Docker->>Transcript: stream and stderr, runner seam Transcript-->>Ledger: parse turns and usage Ledger-->>Runner: agent.reply and tool records Docker-->>Backend: process return code Backend->>Docker: stop on abnormal exit, agent_backends.py 1010 to 1019 Backend->>Fabric: release or cancel, agent_backends.py 1027 to 1045 Backend-->>Runner: AgentRunResult, agent_backends.py 1049 to 1061 Runner->>Runner: agent.handoff_written, run_cell.py 1156 to 1193 Runner->>Runner: agent.pass_ended, run_cell.py 1463 to 1500
7. The records
The canonical records owned here are the pass result, lifecycle, prompt injections, transcripts, stderr, container proof, handoff validation and the prompt retry metadata. The cell manifest is written by the runner while this mechanism supplies its Fabric pass member. Readers are named where the source was verified; later aggregators read the files rather than calling the backend objects.
| Record | Writer | Fields or contents | Readers |
|---|---|---|---|
AgentRunResult | agent_backends.py 1049 to 1061 | process_status, transcript path, error_text, reported_usage, provider identity, model provenance, argv and lifecycle | run_cell.py 2542 to 2547 and result accounting after the pass; the runner then feeds scoring and metrics. |
fabric.passes.<k> in cell_manifest.json | run_cell.py 2522 to 2532 through the backend record callback | Admission and lease identifiers, states, reason, timestamps, waiting seconds and rounds, events, model provenance, context window, process and release durations | The live interface and timing readers use the manifest; the runner reads it on resume at run_cell.py 2534 to 2535. |
| Prompt injection record | run_cell.py 2490 to 2496 and per-pass result assembly | sys_prompt, task_brief and harness_overhead categories with character counts | Metrics and the campaign ledger read the per-pass injection information. |
| Raw transcript | The injected runner called at run_cell.py 2517 to 2518 or the Fabric child path | JSON lines of agent stream events, including provider turns and terminal data | Handoff capture and the runner’s usage and error readers; timing and ledger readers process the stream. |
| Stderr transcript | The injected runner and backend path at run_cell.py 2499 to 2500 and agent_backends.py 1049 to 1050 | Process error text | agent_backends.py error finishers and the runner’s failure record. |
| Claude CID file | agent_backends.py 1122 to 1138 through Docker --cidfile at 473 to 476 | Container identifier for the pass | stop_container_from_cidfile at agent_backends.py 516 to 542 on abnormal cleanup, and evidence readers. |
| Codex profile | agent_backends.py 396 to 406 and 1078 to 1082 | Rendered non-secret provider URL and model alias in cell-local config.toml | Codex inside the measured container reads it through CODEX_HOME. |
| Handoff record | Agent writes workspace HANDOFF.json; run_cell.py captures it at 1156 to 1193 | Structured summary, changed files, interfaces, identifiers, conventions, data, walkthrough, settings and not-done notes | handoff_lib.py 26 to 32 and 160 to 181; the runner stores validation problems as evidence. |
| Handoff summary | handoff_lib.py 183 to 190 | Compact status string | Runner reporting and operator-facing summaries. |
| Controller record | Fabric client calls made through agent_backends.py 799, 977, 1027 and 1042 | Admission, claim, heartbeat, release or cancellation state | The controller, its event collector and the lifecycle callback. |
The record lineage is hand-authored for this plan row. It is not a flow-model rendering. The diagram cites the writer and reader boundaries above and follows the prompt into the child, then into the per-cell and campaign stores.
flowchart LR Template[Prompt template\nrun_cell.py 81 to 104] --> Prompt[Composed prompt\nrun_cell.py 2472 to 2496] Task[Task package\nrun_cell.py 1560 to 1581] --> Prompt Hint[Artifact hint\nrun_cell.py 869 to 922] --> Prompt Retry[Retry feedback\nrun_cell.py 271 to 313] --> Prompt Prompt --> Backend[Backend argv and environment\nagent_backends.py 975 to 978] Backend --> Stream[Pass transcript and stderr\nrun_cell.py 2499 to 2500] Backend --> FabricRecord[Fabric lifecycle\nrun_cell.py 2522 to 2532] Stream --> TurnLedger[Turn and token ledgers] Stream --> Handoff[HANDOFF.json and validation\nhandoff_lib.py 26 to 181] FabricRecord --> CellManifest[cell_manifest.json] TurnLedger --> CellMetrics[per-cell metrics] CellManifest --> Campaign[Campaign tables and timing aggregates] CellMetrics --> Campaign Handoff --> Campaign
8. The loops and the waits
The pass loop is for k in range(start_pass, max_passes + 1) at run_cell.py lines 2464 to 2470. Its upper bound is the configured maximum pass count, and it stops when the range ends or when cumulative tokens reach the budget before composition. The agent request loop is external to the harness. The flow model records it as agent.request; its bound is the agent tool’s own turn limit and token budget, not a numeric bound in these owning files. The backend does not add another request loop around the child.
Lease acquisition is the unbounded while True at agent_backends.py lines 795 to 827. It stops when client.acquire returns, when retryable_acquire_failure returns no retry reason and the exception is raised, or when SystemExit or KeyboardInterrupt unwinds the loop. A retryable failure sleeps for lease_retry_seconds, read at lines 113 to 123 and normally 60 seconds. The interval is measured into lease_wait_seconds in the sleep finally block at lines 817 to 827. Fabric admission polling is inside FabricClient.acquire, which this chapter invokes but does not own, so its exact polling interval and timeout belong to chapter 03. The client call timeout and controller routes are likewise owned there.
Heartbeat waiting is delegated to HeartbeatTask, started at agent_backends.py lines 977 to 978 and stopped at lines 1020 to 1026. Its interval and failure behavior are implemented by the Fabric client module, not duplicated here. The child runner waits for process completion at agent_backends.py lines 981 to 983. Container stop uses a five-second Docker stop grace period and a sixty-second host subprocess timeout at lines 538 to 541. There is no backend timeout around the agent process in these lines. Release and cancellation are single cleanup calls, with failures converted to process status 75 at lines 1021 to 1045.
9. The guards and refusals
| Guard | Location | Refusal or alteration |
|---|---|---|
| Backend name membership | agent_backends.py 373 to 379 | Rejects an unsupported route with SystemExit. |
| Endpoint name membership | agent_backends.py 708 to 712 | Rejects an unknown Fabric endpoint with configuration error. |
| Container requirement for DGX | run_cell.py 1713 to 1717 | Refuses a Fabric backend without FIVEB_CELL_CONTAINER. |
| Lane presence and membership | run_cell.py 1724 to 1732 | Refuses a missing or unapproved maintenance lane. |
| Conflicting Anthropic credentials | run_cell.py 1734 to 1749 | Refuses DGX Claude before staging so a local cell cannot silently bill an Anthropic account. |
| Exact model catalogue presence | agent_backends.py 726 to 732 | Refuses a named-endpoint alias absent from the live catalogue. |
| Interface and capability compatibility | agent_backends.py 733 to 738 | Refuses an alias not ready for the selected backend protocol or missing required capabilities. |
| Positive retry interval | agent_backends.py 117 to 123 | Replaces zero, negative or invalid retry settings with the 60-second default. |
| Prompt argument length | run_cell.py 2481 to 2489 | Bounds overhead to preserve a runnable single argument. |
| Retry feedback parse and resolution | run_cell.py 300 to 313 and 271 to 285 | Uses generic hidden feedback when the map, check number or criterion cannot be resolved. |
| Hint path filters | run_cell.py 891 to 901 | Excludes the inventory, test trees, non-document suffixes and missing files. |
| Secret exclusion from argv | agent_backends.py 448 to 455 and 497 to 500 | Passes tokens by environment name and never embeds their values in the command. |
| Abnormal container cleanup | agent_backends.py 1009 to 1019 | Attempts to stop the exact CID before releasing the lease. |
| Handoff validation isolation | handoff_lib.py 160 to 181 and runner call after 2549 | Returns validation problems as data rather than changing the cell verdict. |
10. The unhappy paths
Each entry has four parts: the intended step, the design reason, the trigger and record, and the cost. The downstream effects name the driver, scorer, aggregator and interface explicitly. Where a component has no direct call in the verified path, that absence is stated.
| Trigger and code path | Record, status and exit code | Downstream effect |
|---|---|---|
A backend name or Fabric configuration is invalid. select_backend lines 373 to 379 or runner guards lines 1713 to 1749 refuse before the child starts. This works by failing before staging so an accidental route cannot spend tokens. The trigger is an unsupported name, missing image or lane, endpoint error, or conflicting credential, and the record is the raised SystemExit message with no transcript or lease. The cost is no agent tokens and no scorer input. | Driver stops the cell with the guard’s nonzero SystemExit path. No backend result is returned, no new scorer record is made, and no campaign aggregate row is produced for that attempted start. The interface sees the surrounding cell failure or absence rather than a live agent. | |
Model qualification fails. _qualify_selected_endpoint_model lines 715 to 738 is meant to admit only an alias whose endpoint evidence supports the route. It is inside run_pass so a refusal is recorded as an unstarted pass. A missing alias, unready interface or missing capability triggers a FabricConfigurationError; the lifecycle has its initial fields and the runner callback receives it, but no admission or transcript is created. The cost is no model usage and a failed pass setup. | The driver receives a backend exception through run_pass and writes the surrounding failure state. The scorer has no agent output to score. Aggregation can retain lifecycle evidence if the cell manifest was written, while the interface can report an unstarted or failed pass from that manifest. | |
Lease admission remains temporarily unavailable. _acquire_lease lines 768 to 827 is meant to wait rather than call a failed cell. The controller reason is classified as retryable, so the code records a waiting_for_lease event and sleeps. The trigger is a transient controller, maintenance, queue or connectivity refusal; the record is the growing lifecycle with wait rounds and seconds, and no transcript because the agent has not started. The cost is elapsed wall time and no tokens. | The driver remains in the pass loop until a lease arrives or a stop signal occurs. The scorer is not called for an unfinished pass. Timing aggregation can separate lease waiting from agent time, and the interface can display the waiting lifecycle. | |
Lease admission is terminally refused. The same step is meant to distinguish an impossible request from a temporary refusal. retryable_acquire_failure returns no reason and _acquire_lease re-raises. The trigger is invalid authentication, an unknown alias, malformed response or another non-retryable error; the record is the lifecycle callback state and exception, with no agent transcript. The cost is failed setup and no tokens. | The driver records failure and does not score an agent result. The aggregator can consume the failure manifest if present. The interface sees the terminal reason rather than a running lease. | |
The agent exits nonzero. The execution step is meant to preserve the child evidence and release resources. run_pass lines 981 to 989 marks the outcome failed; cleanup at lines 1009 to 1047 stops the container and releases the lease. The trigger is a nonzero child return code; the record is the transcript, stderr, failed AgentRunResult, lifecycle and error text. The cost is the tokens consumed before failure and one failed pass. | The driver can offer a retry notice on the next pass when the surrounding scoring path classifies the failure. The scorer can inspect the transcript and later product state, the aggregator can count usage and failure status, and the interface can show the failed lifecycle and stderr path. | |
The process is interrupted. The execution step is meant to distinguish cancellation from ordinary failure. run_pass catches KeyboardInterrupt and SystemExit at lines 990 to 1000, sets cancelled, records interrupted_at and re-raises. The trigger is the stop signal; cleanup still removes environment, stops the CID container, stops heartbeat and releases or cancels. The record is the partial transcript where available and a cancelled lifecycle. The cost is the work and tokens already spent, without a retry caused by treating cancellation as a test failure. | The driver exits its stop path rather than continuing the pass loop. The scorer does not score a cancelled attempt. Aggregation can preserve partial timings and tokens, and the interface can show cancellation and cleanup state. | |
Heartbeat or release fails. The intended cleanup step is to keep controller state honest. HeartbeatTask.stop and client.release are guarded at lines 1020 to 1036. A Fabric error changes the process status to 75 and records redacted detail, even if the child had returned zero. The cost is a failed infrastructure finalisation after possible model usage. | The driver receives an AgentRunResult with failure status when cleanup completes. The scorer sees an infrastructure failure rather than a clean pass, the aggregator can retain release timing and error data, and the interface can show the failed release or heartbeat. | |
| A queued admission exists but no lease is returned. Cleanup at lines 1037 to 1045 is meant to prevent abandoned scheduler state. Polling, claim or local startup triggers this branch; it records a cancellation attempt and no agent transcript. The cost is failed setup and no model tokens, with a possible status 75 if cancellation itself fails. | The driver records the failed pass. The scorer and its downstream aggregator have no agent output, while the interface can see the admission and cancellation lifecycle. | |
The container cannot be stopped. The abnormal cleanup step is meant to prevent orphaned compute. stop_container_from_cidfile lines 516 to 542 returns false for a missing CID, an exited container, Docker failure or timeout and never raises. The trigger leaves a cleanup event only when a stop was issued, so the missing-stop case may leave no new file. The cost is possible orphaned compute and operator cleanup, not a new agent result. | The driver continues release cleanup rather than hiding the primary failure. The scorer and aggregator retain the underlying pass outcome. The interface can show the failed pass, but a failed stop is not separately guaranteed as a durable record. | |
| The prompt input or retry map is unusable. Composition is meant to keep the prompt deterministic and within the process argument limit. A missing task file yields an empty contribution, an unreadable retry map selects generic feedback, and an oversized prompt bounds overhead at lines 2481 to 2489. The records are the injection lengths and retry mode, with no separate prompt-error file. The cost is reduced context or an empty task contribution, rather than a hidden substitution. | The driver may still invoke the agent with the composed string unless another surrounding guard refuses. The scorer and aggregator see whatever pass evidence results. The interface can show injection lengths and retry metadata, but not an independent prompt failure status. | |
Handoff JSON is missing, malformed or schema-invalid. Handoff capture is meant to preserve evidence without changing the verdict. load lines 26 to 32 returns None on read or parse failure, and validate lines 160 to 181 returns validation problems. The trigger leaves a handoff validation record with problems or absence and does not alter the backend process status. The cost is loss of structured explanation, not loss of the agent transcript. | The driver and scorer continue according to the process and product evidence. The aggregator can mark the handoff evidence as incomplete, and the interface can display validation problems if it reads the record. |
11. The metrics produced or fed
The backend feeds process and infrastructure facts into the runner, while the prompt path feeds injection and retry facts. AgentRunResult.process_status becomes the pass return code used by the runner at run_cell.py 2542 to 2547. raw_transcript_path and stderr_path identify the files later parsed for turns, errors and timing. For Codex, reported_usage comes from the summed turn.completed events in parse_codex_usage lines 545 to 605. Its required input and output token fields become usage totals, while cached input and reasoning output remain nullable when the provider omits them. For Claude, the backend deliberately returns None at lines 1157 to 1160 because the registered Claude ledger owns those usage columns and a second provider record would double count.
The lifecycle fields written through _record_fabric at run_cell.py lines 2522 to 2532 feed timing and interface views. lease_wait_seconds is advisory for agent duration but is subtracted from cell seconds to success by the runner. lease_wait_rounds, admission_seconds, agent_process_seconds and release_seconds become named timing inputs. The claim event fields exact_model, revision, runtime version, engine, quantization, container digest, request preset and reasoning policy become provenance fields in the manifest and later campaign records. The context window is promoted to cell_manifest.json at lines 2524 to 2531 and feeds the context and compaction interpretation.
The injection categories written at run_cell.py lines 2490 to 2496 become per-pass advisory columns or ledger rows: system prompt characters, task brief characters and harness-overhead characters. They describe what was composed, not what the agent actually read. The prompt variant record from compose_artifact_hint lines 879 to 921 supplies variant, applied flag, file list, markers and character count. Retry metadata assembled after the pass identifies whether feedback was available, which mode was used, and whether the notice was visible, environment, generic hidden or criterion-specific hidden. These are explanatory factors, not outcomes.
The transcript is the source for the turn and token ledgers. The flow-model nodes agent.reply, agent.tool_call.read, agent.tool_call.search, agent.tool_call.edit, agent.tool_call.test_run, agent.tool_call.other, agent.tool_result, agent.compaction and agent.notice are classified from stream events by the ledger and timing readers named in the model. A handoff validation problem is advisory evidence. It does not become a scorer verdict, because the runner comment at the capture call says that a missing or malformed handoff never changes the cell verdict.
12. The tests
The verified source includes test seams rather than a verified test inventory for this mechanism. run_cell.py lines 316 to 318 states that the real Claude executable is never called in tests and that the invocation seam is monkeypatched. The source also contains the fixed-clock rationale for admission timing at agent_backends.py lines 958 to 967 and the explicit lease-wait accounting rationale at lines 821 to 824. Those comments show intended test seams, but no test file was verified in the allowed harness and interface paths for every branch in section 10.
The mapping of tests to exit paths was not made. In particular, no verified test file was established here for terminal catalogue refusal, each retryable lease classification, process nonzero cleanup, interruption and re-raise, heartbeat failure, release failure, admission cancellation, failed container stop, prompt length bounding, each retry notice form, or malformed handoff validation. The only checked statement for section 10 is therefore that a complete trigger-to-test mapping is absent. This is an absence of verification, not a claim that no tests exist elsewhere in the repository.
13. The dated incidents that shaped the code
The comments in the owning files carry dated reasons for guards that might otherwise look arbitrary. At run_cell.py lines 81 to 103, the single prompt and v006 handoff and documentation amendments are dated 2026-09-02 and 2026-09-03. They explain why there is one prompt of record, why the successor handoff is mentioned, and why the closing documentation instruction is conditional. The prompt variant rationale at lines 110 to 126 records the 2026-09-01 ruling that an artifact hint is a measured addition rather than a new prompt.
The retry notice history at run_cell.py lines 178 to 198 records the 2026-07-26 exposure incident. Earlier wording disclosed hidden checks and encouraged searching for them. The replacement notices use the task description and visible evidence instead. Lines 206 to 225 record the 2026-08-17 environment-dependent path incident, in which a container-visible absolute path passed locally and failed outside. Lines 235 to 253 record the 2026-09-01 ruling for criterion-specific hidden feedback and the generic fallback when the frozen map cannot resolve a criterion.
The container and lease safeguards have corresponding history. agent_backends.py lines 473 to 475 date the CID-file proof and abnormal-stop safeguard to the 2026-08-29 transparency audit. Lines 519 to 525 describe the 2026-08-29 run set 039 stop in which a killed Docker client left its container computing. Lines 991 to 996 record the 2026-08-29 correction that SystemExit must take the cancelled path rather than be swallowed as failure. The Claude backend comments at lines 1107 to 1112 record the 2026-08-27 ruling that the Anthropic runner is the measured local route and that exactly one authentication source is allowed. The request preset and reasoning provenance comment at lines 653 to 658 records the 2026-09-05 Fabric change that made those claim fields necessary because the runner cannot select its own reasoning setting.
The prompt length guard at run_cell.py lines 2481 to 2489 is also historical. Its comment says an overlong single argument caused earlier cells to fail mid-loop and that trimming the retry notice is worse than losing later attempts. The lease wait comments at agent_backends.py lines 774 to 793 preserve the operator ruling that transient controller unavailability is not a cell failure, while stop signals must unwind immediately. These dates and reasons explain why the code records waiting and cancellation separately from ordinary failure.
14. The weakest claim, what was not checked, and the token line
The weakest claim is the downstream column mapping from every lifecycle and injection field into campaign tables, because the owning files verified here write the source records while later aggregators and interface readers spread their consumption across other chapters. The exact tests covering each unhappy exit path were not checked, and the mapping of tests to exit paths was not made. The controller’s polling timeout and heartbeat interval are delegated to the Fabric client and are not restated as facts beyond the calls verified here.
Model: openai/gpt-5.6-luna via OpenRouter; tokens: see the job ledger