Written 2026-09-12 by Claude Fable 5.1 to the standard set in 5. Experiment/11. Detailed Design/specifications/C0-detailed-design-specification.md and shown by the reviewed sample chapter 06-execution-of-a-cell.md. Every line number below was read from the source at commit 6e2ee2c2b; the four owning scripts have the same Git blob hashes at that commit as at commit b69ae5977, the commit the flow model 5. Experiment/11. Detailed Design/flow-model/flow_model.v001.json was verified against, so the two agree line for line. The local-model draft operations/detailed-design/drafts/chapter_09.md was the lead for the inventories and nothing else; section 14.1 lists what it got wrong. The flow model’s chapter 09 subgraph (nine nodes, seventeen edges, all witnesses verified) is the source of the state diagram of section 5, and every node identifier is named there so the renderer can regenerate it.
Owning files. 5. Experiment/1. Harness/scripts/timing_lib.py (1,681 lines, 82,127 bytes), the stage timer and the transcript readers; 5. Experiment/1. Harness/scripts/ledger_lib.py (539 lines, 22,560 bytes), the token ledger builder; 5. Experiment/1. Harness/scripts/regenerate_token_ledgers.py (150 lines, 5,671 bytes), the offline ledger rebuild; 5. Experiment/1. Harness/scripts/backfill_turn_ledgers.py (215 lines, 9,878 bytes), the offline turn-ledger backfill; and the two normative documents 5. Experiment/6. Metrics/token_ledger_spec.md (22,210 bytes) and 5. Experiment/6. Metrics/generation_ledger_spec.md (15,339 bytes). Files this chapter reads into but does not own: run_cell.py (the cell runner, which calls every writer described here, chapter 06), score_cell.py (the scorer, which appends to the timing file, chapter 08), run_batch.py (the batch driver, which appends its own view of each cell, chapter 02), aggregate_metrics.py and aggregate_timings.py (the two aggregators that turn these records into campaign tables, chapter 10), agent_backends.py (the lease lifecycle the timing record copies, chapter 03), and 10. User Interface/server.py (the results interface, chapter 11).
1. What the two records are and why they exist
A cell is one measured attempt: one coding agent, one maintenance task, one product build, one repetition seed, run at most five times within one cumulative token budget. Everything the experiment says about a cell’s cost comes from two records this chapter owns. The token ledger, the file token_ledger.jsonl inside the cell’s folder, says what the agent’s conversation with the model cost, one row per event of the conversation, with every token attributed to a category such as reading code, reading documentation, editing code, or running tests. The timing record, the file timings.jsonl in the same folder, says how long each separate activity of the cell took, one line per finished activity, so that a forty-minute cell can be read as thirty minutes of waiting for the model, six minutes of the agent running the product’s test suite, and four minutes of the harness checking the result, rather than as one undifferentiated number.
Three smaller records travel with them. The turn ledger, one file per pass named transcripts/pass<k>.turns.jsonl, is one line per request the agent sent to the model, with what the request waited, what it streamed, how much conversation it carried, and every tool call the reply then made. The live marker, timing_current.json, names the activity in progress and is deleted when the cell finishes, so that its survival is the mark of a crash. The provider usage record, provider_usage.jsonl, holds the usage figures the Codex back end reports, for the one route whose transcripts the ledger algorithm cannot read.
None of these writers can fail the cell. Every write in timing_lib.py is wrapped so that a full disk, a read-only mount, or a value that will not serialise loses the line and nothing else (module docstring, lines 28 to 33). The reasoning the comments give is that a cell costs real money and hours of machine time, and a diagnostic recorder that could destroy a measured result in exchange for a diagnostic convenience would be a bad trade at any price. The ledger builder is the one exception: it raises on a transcript that will not parse, and the runner does not wrap the call, so a corrupt transcript crashes the cell (section 10).
The records are written by the cell runner while the cell runs (chapter 06 is the caller; the per-pass sequence is section 4 here), appended to by the scorer and the batch driver afterwards, and read by the aggregators, the report generator, the results interface and two offline rebuild tools. The token ledger is the source of the campaign’s primary metric, tokens_to_success, which the aggregator recomputes from the ledger rows rather than copying from the runner’s own sums (section 11).
2. The reader’s map of the owning files
2.1 timing_lib.py
Six regions. Lines 1 to 41 are the module docstring, which states the four design rules the file follows: what it records, why a separate file, why it can never fail the cell, and what the live marker is. Lines 42 to 64 are the imports and the constants: the schema tag TIMING_SCHEMA equal to 5b.cell_timings.v001 (line 54) and the five phase names setup, pass, finalize, scoring and batch (lines 59 to 64), kept as a fixed list because the interface groups by phase and a free-form string would let a typo create a silent sixth category. Lines 67 to 304 are the timer: _Span (line 71), the record of one activity in progress; StageTimer (line 88), with _append (118), _mark_live (125), stage (147), open (187), close (205), record (229) and clear_live (270); and NullTimer (278), a timer that records nothing, for callers with nowhere to write. Lines 307 to 418 roll the lines up: the container list CONTAINER_STAGES (333), read_timings (337) and summarize (352). Lines 421 to 1145 read a finished pass back out of its transcript: _epoch (421), _stream_events (433), transcript_timing (445), request_summary (568), request_records (606), the test-command constants (758 to 763), _command_runs_tests (766), _tool_result_text (793), turn_ledger (818) and ledger_summary (1054). Lines 1147 to 1681 are the additions of 2026-09-03: the restated new-context arithmetic new_context_summary (1274) and new_context_tokens_total (1329), harness_overhead_seconds (1366), the turn-ledger file functions write_turn_ledger (1399) and read_turn_ledger (1434), the column list PASS_TIMING_COLUMNS (1466), the machine-wait stage list (1483), _median (1486), pass_timing_rows (1496) and cell_timing_from_passes (1633).
2.2 ledger_lib.py
Five regions. Lines 1 to 42 are the module docstring, which names the specification the file implements (token_ledger_spec.md, sections 5 to 7) and records the two parser corrections, v002 of 2026-08-28 and v003 of 2026-09-02 (section 13). Lines 44 to 64 are the constants: the parser version 5b.ledger_parser.v003 (44), the ten run categories (55 to 60), and the command and tool sets the classifier consults (61 to 64). Lines 67 to 209 are the classifier: _lap_artifact_id (67), _is_comment_only_edit (84), the four shell-command helpers (94 to 131), _classify_bash (134), workspace_relative (150) and classify_event (169), which is the single source of the specification’s section 7 decision table. Lines 212 to 228 are load_lap_manifest, the reader of the build’s artefact inventory. Lines 231 to 417 are the stream readers: _est (231), _distribute (235), _new_row (253), _read_stream (267), _normalize_stream (276), _reported_totals (380) and reported_usage_totals (415). Lines 420 to 539 are build_ledger, the deterministic algorithm, with its three steps marked in comments: step A at line 431, step B at 441, step C at 526.
2.3 regenerate_token_ledgers.py
Three regions. Lines 1 to 23 are the docstring, imports and constants, including the set of back ends whose ledgers can be rebuilt, CLAUDE_BACKENDS equal to claude and dgx_claude (line 21). Lines 26 to 110 are _category_totals (26), _pass_number (36) and rebuild_cell (43), which rebuilds one cell and returns whether anything changed. Lines 112 to 150 are rebuild_run_set (112) and main (125).
2.4 backfill_turn_ledgers.py
Three regions. Lines 1 to 56 are the docstring, which states what a request ledger is, why the script exists, that it costs the agent nothing, and where it may be run, and the imports. Lines 59 to 109 are the three helpers batch_is_running (59), visible_test_command (79) and transcripts_of (97). Lines 112 to 215 are backfill (112) and main (160).
3. The inputs
3.1 Command-line arguments
The two library modules have no parser. regenerate_token_ledgers.py parses at lines 126 to 130: --run-set (repeatable, required; a folder path or a run-set name resolved under 5. Run Sets, lines 133 to 138) and --check (report changes without writing them). backfill_turn_ledgers.py parses at lines 161 to 178: --base (the 5. Experiment folder), --run-set (repeatable, prefix match), --overwrite, --dry-run, --force (run even while a batch says it is running) and --aggregate (rebuild the campaign timing tables afterwards).
3.2 Environment variables
None of the four files reads an environment variable. backfill_turn_ledgers.py uses the interpreter path sys.executable (line 202) to start the timing aggregator as a child process. The runner’s environment variables that shape what these records contain are chapter 06’s (section 3.2 there).
3.3 Files read
| File | Read at | What it decides |
|---|---|---|
cells/<cell>/transcripts/pass<k>.stream.jsonl | ledger_lib._read_stream 267 (through build_ledger 425 and reported_usage_totals 417); timing_lib._stream_events 433 (through transcript_timing 490, write_turn_ledger 1417, pass_timing_rows 1562) | The whole of the ledger and the turn ledger; the transcript is the only source of both |
<variant root>/lap_artifact_manifest.json | ledger_lib.load_lap_manifest 212 to 228 | Which paths are documentation artefacts (lap_files) and which are source files carrying trace anchors (trace_anchor_sites united with the legacy probe_comment_sites); a missing file yields empty sets, so every path classifies as ordinary code |
cells/<cell>/timings.jsonl | timing_lib.read_timings 337 (through summarize 361, harness_overhead_seconds 1381, pass_timing_rows 1530) | The roll-up and the per-attempt rows |
cells/<cell>/transcripts/pass<k>.turns.jsonl | timing_lib.read_turn_ledger 1434 (through pass_timing_rows 1559 and the runner at run_cell.py 2591) | The per-request figures; when absent, pass_timing_rows rebuilds them from the transcript and labels the row recomputed (1560 to 1566) |
cells/<cell>/metrics.json | regenerate_token_ledgers.rebuild_cell 45 to 46 and rebuild_run_set 117 | The back end, the per-pass injection sizes, and the record to rewrite |
cells/<cell>/staged_snapshot/lap_artifact_manifest.json | regenerate_token_ledgers.rebuild_cell 56 | The artefact inventory as the cell saw it, so a rebuild classifies against the same manifest the run did |
cells/<cell>/cell_manifest.json | backfill_turn_ledgers.visible_test_command 79 to 94 | The registered visible test command, so a shell command in the transcript can be recognised as a run of the product’s suite |
5. Run Sets/active_batch.md | backfill_turn_ledgers.batch_is_running 59 to 77 | Whether a batch says it is running, in which case the backfill refuses unless forced |
3.4 Network
None. Every function here reads files the runner already left on disk. The backfill starts one child process, aggregate_timings.py, at lines 199 to 202.
4. The happy path in order
The order is the order of one pass of the cell runner, then the runner’s finalisation, then the two later appenders, then the two offline tools. The runner’s line numbers are chapter 06’s territory and are cited here because the flow model places these nodes in chapter 09.
4.1 The timer is constructed, run_cell.py line 1806
StageTimer (timing_lib.py line 88) is built once per cell with the cell folder, the run identifier and the cell identifier (__init__, lines 107 to 116). It keeps a stack of open activities (_stack, line 115) so that each finished line can carry the name of the activity that contains it and its depth. Two clocks are used on purpose: the duration is measured on the monotonic clock, which is immune to the system clock being adjusted under a long cell, and the start and end stamps come from the wall clock so that a human can line an activity up against the transcript and the container logs (docstring, lines 97 to 100).
From here to the end of the cell the runner names its activities. The setup phase has stage_working_copy (2036), hash_before (2062), scan_task_brief (2088), inject_seeded_defect (2096 to 2132, through open and close because the block has early exits), inject_feature_removal (2137 to 2158), staging_gate (2171), initialize_workspace_git (2191), snapshot_staged_tree (2197) and snapshot_manifests (2204 to 2354). The pass phase has the container pass (2466), agent_invocation (2503, with the back end, whether it ran in a container and the prompt length as attributes), capture_handoff (2553), build_token_ledger (2647), repo_escape_check (2688), visible_tests (2779), visible_tests_host_confirm (2814), make_patch (2826) and hidden_checks (2850). The finalisation phase has remove_workspace_git (2918) and hash_after (2922). Each stage block writes one line when it ends, with ok false and the exception re-raised unchanged when the block died (lines 147 to 185), so a stage that failed after twelve minutes still leaves the seconds it burned.
4.2 The pass span and the agent span, lines 2464 to 2547
At the top of each iteration k the previous pass span is closed and a new one opened (2465 to 2466). The agent’s turn is opened at 2503 and closed at 2547 with the return code as an attribute and ok equal to whether it was zero. Everything between is chapter 06’s business; what matters here is that the span’s start stamp, _agent_span.started_iso, is later used to place the agent’s own report of itself on the clock (section 4.5).
4.3 The lease lifecycle placed on the clock, lines 2564 to 2574 (flow-model node timing.lease_recorded)
On a Fabric route the back end returns its lifecycle dictionary (chapter 03, agent_backends.py run_pass), and the runner writes four lines through record (timing_lib.py 229 to 268): lease_wait with the number of rounds as an attribute (2566 to 2568), lease_admission (2569), agent_process (2571) and lease_release (2573). record writes a duration something else measured; its lines carry measured false so that an analysis can keep the harness’s own stopwatch readings apart from figures reported by the thing being measured (docstring, 233 to 240). On the Anthropic route there is no lifecycle and the four lines are absent, which is the witness the flow model uses to tell the two routes apart (edge cell.pass.capturing->ledger.turn_ledger_written).
4.4 The transcript split and the turn ledger, lines 2575 to 2598 (node ledger.turn_ledger_written)
transcript_timing (timing_lib.py 445 to 565) reads the pass transcript and returns two accounts of the same minutes. The first is what the agent’s command-line tool reports about itself in its terminal result event: the session length, the part spent on model requests, the time to the first token and the number of turns (lines 495 to 509). The second is the same split recomputed from the event timestamps rather than taken on trust: the gap from a tool result going back to the next assistant message is the model thinking, and the gap from an assistant message to the tool result coming back is the agent’s own command running (lines 513 to 553, with tool calls paired to results by identifier and not by adjacency, 531 to 550). It then calls request_records (606 to 755) for one record per model request and folds the totals in (555 to 564).
The turn ledger is then written at 2586 to 2591: write_turn_ledger (1399 to 1431) parses the transcript again through turn_ledger (818 to 1051) and writes one JSON object per request to transcripts/pass<k>.turns.jsonl. turn_ledger groups the assistant events of one reply by their shared message identifier (822 to 829 explain that adjacency was tried first and turned 45 real replies into 58), splits each round trip three ways (the wait from the moment the agent could send to the first block of the reply, the streaming of the rest, and the two added), and pairs every tool-use block to its result by identifier, keeping a call that never got a result with its end fields null (docstring, 851 to 860). A shell command is marked runs_tests when its text contains one of five runner names or the cell’s registered visible test command (_command_runs_tests, 766 to 790), and sleeps when it matches the sleep pattern (763). The whole write is inside a contextlib.suppress(Exception) block at 2588, so a failure loses the ledger and never the cell; the rows are read straight back at 2591 and reduced by ledger_summary (1054 to 1145) at 2597.
4.5 The agent’s report placed on the clock, lines 2607 to 2634 (node timing.agent_report_recorded)
Three more record lines follow, each carrying derived_from equal to transcript and the agent span’s own start and end as the line’s start and end: agent_model_api with the time to first token, the turn count and the merged request summary (2607 to 2621), agent_own_commands with the per-tool seconds, the tool-call count, the test-run count and seconds, and the sleep seconds (2622 to 2630), and agent_model_gaps (2631 to 2634). The comment at 2600 to 2606 records why the start and end are passed: until 2026-09-03 these lines had no start, and the results interface could list them but had nowhere to draw them, so the operator saw one long bar labelled the agent’s turn and nothing inside it. Passing the stretch lets the figure be drawn inside the bar it belongs to while staying marked as not measured, because placing a figure does not turn a report into a stopwatch reading (timing_lib.py 241 to 251).
4.6 The token ledger built and appended, lines 2645 to 2653 (node ledger.token_ledger_built)
For the two Claude-shaped back ends, build_ledger (ledger_lib.py 420 to 539) is called under the stage build_token_ledger (2647 to 2650) and the rows it returns are appended to token_ledger.jsonl (2651 to 2653). The algorithm is the specification’s section 5, in three steps.
Step A (431 to 439) emits one harness_injection row per injected text, in the order the runner passes them (the system prompt, the task brief and the harness overhead, each with its category and its character count), with all four token fields zero and a weight equal to the text’s estimated size, one token per four characters (_est, 231).
Step B (441 to 524) walks the normalised transcript in order. The normalisation, _normalize_stream (276 to 377), collapses the several assistant events that one reply arrives in, all sharing a message identifier, into one logical message carrying the largest usage snapshot seen for that identifier (the comment at 329 to 338 explains that streamed usage is cumulative per request and never decreases, so the per-field maximum is the request’s whole usage, placed on the first fragment so that the request’s input is attributed to what preceded it and later fragments carry zero). When no logical message carries any output but the terminal result event reports a positive output total, the shape of the locally served Anthropic-compatible stream, that total is distributed across the messages by their estimated size and each is marked _synthetic_output (365 to 376). For each logical assistant message the input side (input, cache-creation and cache-read tokens) is attributed to what the model read to produce it: the injection rows for the first message, and for later messages the tool-call rows of the previous message whose results came back, weighted by the size of each result (467 to 481). The output side is split across the message’s content blocks by estimated size (483 to 490); a tool-use block becomes a tool_call row classified by classify_event (169 to 209), and any other block becomes an assistant_message row of category reasoning_out (511 to 518). The attribution method is reported only when the message has a single block and its output was not synthesised, and estimated_chars4 otherwise (488 to 489, 496). Two notes are set here: an unknown tool gets unknown_tool_fallback:<name> and the method estimated (499 to 500, rule T1 of the specification), and a shell command classified as a code read only because nothing else matched gets bash_fallback:<word> (501 to 504, rule B5).
The classifier’s rules, in classify_event (169 to 209) and _classify_bash (134 to 147): a text block is reasoning_out (178); a Read of a manifest-listed documentation file is lap_read with its artefact identifier, and of anything else code_read (181 to 186); Grep, Glob, LS and WebSearch are search_nav (188 to 189); an Edit, Write or NotebookEdit of a documentation file is doc_maintenance (191 to 195), of an anchor-bearing source file where only comments or docstrings changed is doc_maintenance with a trace_comments: identifier (196 to 199, the E2 rule, decided by _is_comment_only_edit 84 to 91 through strip_lap.strip_python_source), and otherwise code_edit (200); a shell command is test_exec when it is a pytest or unittest invocation (120 to 131, 138 to 139), search_nav for the search words (140 to 141), lap_read or code_read for the read words and sed -n (142 to 146), and code_read for everything else (147). Paths are normalised to the workspace-relative form the manifest uses by cutting at the last /workspace/ marker (workspace_relative, 150 to 166).
4.7 The reconciliation remainder, lines 526 to 538 (node ledger.reconciliation_remainder)
Step C computes the reported per-pass usage, _reported_totals (380 to 412), which sums the normalised assistant messages field by field and then takes the per-field maximum against the last result event’s usage, so that a larger terminal figure is authoritative and an error result’s explicit zeros never erase usage already present (docstring, 381 to 386). The remainder per field is the reported figure minus the sum of the rows (528 to 529). When any field’s remainder is not zero, one more assistant_message row of category reasoning_out carries the whole difference, with notes equal to reconciliation_remainder pass <k> (531 to 537). The rows therefore always sum exactly to the reported usage, by construction, and the remainder row is where unattributable output such as thinking tokens lands. The notes prefix is the machine-readable convention every consumer uses to find the row (specification section 8; the aggregator at aggregate_metrics.py 478 to 481 and 257 to 258).
4.8 The reconciliation check, lines 2654 to 2661 (node ledger.reconciliation_check)
The pass spend is the sum of input, output and cache-creation tokens over the rows (2654 to 2655), which is the cache policy of the specification’s section 9: cache reads are recorded but not spent. The runner then reads the reported usage a second time through _reported_usage (1401 to 1409, which delegates to reported_usage_totals, ledger_lib.py 415 to 417) and compares it field by field with the row sums; a gap above half a percent of the reported figure or fifty tokens, whichever is larger, sets ledger_reconciliation_warning on the result (2656 to 2661). Because step C has already made the rows sum exactly, this warning can only fire when the remainder row itself is large; it is a measure of how much of the pass could not be attributed, not of a broken ledger. The flag reaches metrics.json at 2976 to 2977.
4.9 The Codex usage record, lines 2672 to 2680 (node ledger.provider_usage_appended)
For the Codex back end the reported usage is appended as one line to provider_usage.jsonl with the run and cell identifiers, the pass, the provider name and the model alias (2673 to 2677), and the pass spend is the provider’s total_tokens or zero when the usage is absent (2678). The comment at 2641 to 2644 states the rule: the Codex stream is preserved first and parsed into a separate record that is never passed through the Claude ledger algorithm. A Codex pass with no usage at all is halted by chapter 06’s guard before it can be recorded as zero (run_cell.py 2757 to 2772). The cumulative spend advances at 2679 for both routes.
4.10 The roll-up, the category totals and the live marker, lines 3013 to 3023 (nodes timing.summary_rolled_up and timing.live_marker_cleared)
At finalisation the runner reads its own ledger back and sums the spend per category into category_totals (_category_totals_from_ledger, 3057 to 3062, at 3013), reads its own timing file back and rolls it up into stage_timings (_timing_summary, 3032 to 3045, at 3020), writes metrics.json (3022) and removes the live marker (clear_live, timing_lib.py 270 to 275, at 3023). summarize (352 to 418) applies two rules that the comment at 307 to 325 says matter because the naive sum is wrong in two different ways: an activity that contains others (the container list at 333: pass and the four driver views) is written but never added to a total, and a line whose measured is false is reported under breakdown_seconds rather than added, because it describes seconds already counted from the inside. The roll-up carries accounted_seconds, harness_overhead_seconds (computed by harness_overhead_seconds, 1366 to 1396, which leaves out the containers, the agent’s turn and every reported line), by_stage, by_phase, breakdown_seconds and per_pass. It is deliberately not written as another line in the file, because every re-score would append it again (comment, 386 to 392). _timing_summary returns None when nothing was recorded, so a record written by a harness that could not write timings looks as it did before rather than carrying an empty structure (3033 to 3038).
4.11 The scorer and the driver appending to the same file after the runner
The scorer builds its own StageTimer on the same cell folder (score_cell.py 1277) and appends its stages under the scoring phase: preserve_prior_scoring (1280), score_visible_suite (1309), score_full_suite (1313), score_hidden_checks (1323) and score_write_ledger (1330), then rewrites the stage_timings key of metrics.json from the now longer file (_refresh_stage_timings, 1403 to 1424, called at 1398) and clears the live marker (1399). The comment at 1273 to 1276 gives the reason one file holds both programs’ lines: scoring is a separate process, which is why the phase is recorded on every line rather than inferred from position. The evidence capture that follows scoring is timed by a second timer with no live marker (capture_evidence, 1440 to 1444). The batch driver builds a third timer with no live marker (run_batch.py 461 to 462) and appends its outside view of the two child programs under the batch phase: driver_run_cell (491), driver_score_cell (500), driver_resume_run_cell (529) and driver_resume_score_cell (536). Those four are containers (CONTAINER_STAGES, 333 to 334); their value is the difference between the driver’s reading and the child’s own total, which is the cost of starting an interpreter and importing the harness (comment, 327 to 332).
4.12 The offline rebuild of the token ledger, regenerate_token_ledgers.py
rebuild_cell (43 to 110) reads metrics.json and returns False at once for a cell whose back end is not Claude-shaped (46 to 47). It lists the pass transcripts by number (50 to 54), loads the artefact manifest from the staged snapshot (57), and for each transcript rebuilds the injection list from the per-pass record’s categories and character counts, using a placeholder text of the recorded length because the builder uses only the category and the size (66 to 69, and the docstring at 4 to 6), calls build_ledger (70 to 73), and checks that the rebuilt rows sum to the reported usage in every field, raising ValueError otherwise (74 to 79). It recomputes the per-pass spend, the cumulative spend and the per-pass category totals into the metrics record (81 to 86), sets the cell’s cumulative_spend and category_totals and clears ledger_reconciliation_warning (89 to 91), and writes both files atomically through temporary files and replace only when the text differs (93 to 109). main (125 to 147) exits non-zero in --check mode when a rebuild would change something (146), so the check can be used as an assertion that ledgers on disk are the ones the current parser would produce.
4.13 The offline backfill of the turn ledger, backfill_turn_ledgers.py
main (160 to 209) refuses to start while active_batch.md says a batch is running unless --force or --dry-run is given, returning 2 (180 to 186), because the script writes inside 5. Run Sets and would race that batch’s closeout, which reads the same folders (docstring, 28 to 32). backfill (112 to 158) walks every run set and cell, and for each pass transcript writes a ledger only where none exists, through write_turn_ledger (146 to 147), counting cells, attempts, ledgers written, ledgers skipped, transcripts holding no request and requests written. A transcript that holds no model request at all is counted as empty rather than lost (151 to 156). With --aggregate the campaign timing tables are rebuilt by running aggregate_timings.py as a separate process, so that a failure there cannot lose the ledgers already written (197 to 209).
5. The state machine of the per-pass record chain
The states are the nine nodes of the flow model’s chapter 09 subgraph and the seven chapter 06 nodes they connect to; every node label below is the model’s identifier, and every transition carries the line range and the witness the model records for it. A state is entered when the named code runs; the witness is the record condition a later reader can check to know it ran. The diagram is the model’s subgraph rendered as Mermaid; a renderer can regenerate it from the nodes whose chapter is 09 and the edges with an endpoint among them.
stateDiagram-v2 state "cell.pass.capturing (06)" as CAP state "timing.lease_recorded" as LEASE state "ledger.turn_ledger_written" as TURN state "timing.agent_report_recorded" as REPORT state "cell.pass.capturing.command_logged (06)" as CMD state "cell.pass.ledgering (06)" as LEDG state "ledger.token_ledger_built" as BUILT state "ledger.reconciliation_remainder" as REM state "ledger.reconciliation_check" as CHECK state "ledger.provider_usage_appended" as PROV state "cell.pass.guarding (06)" as GUARD state "cell.finalising.hash_after (06)" as HASH state "timing.summary_rolled_up" as ROLL state "cell.finalising.metrics_written (06)" as METW state "timing.live_marker_cleared" as LIVE state "cell.finalising.manifests_finalised (06)" as MAN CAP --> LEASE : Fabric route, run_cell 2564..2574 (stage lease_wait exists) CAP --> TURN : Anthropic route, 2564..2589 (no lease_wait line, backend claude) LEASE --> TURN : 2575..2589 (lease_wait.measured false; admission, process, release lines exist) TURN --> REPORT : 2589..2600 (turn ledger k exists; agent_model_api line exists) TURN --> REPORT : unhappy, 2588..2591 (turn ledger k absent; agent_model_api line exists) REPORT --> CMD : 2636 (agent_model_gaps line; command_log row) LEDG --> BUILT : Claude-shaped back end, 2645..2651 (stage build_token_ledger; backend in claude, dgx_claude) LEDG --> PROV : Codex back end, 2670..2680 (provider_usage row; backend dgx_codex) BUILT --> REM : ledger_lib 527..538 (build_token_ledger.ok; at least 3 rows for pass k) REM --> CHECK : rows sum exactly, ledger_lib 531 (no reconciliation_remainder note in pass k) REM --> CHECK : unhappy, remainder row appended, 531..537 (note matches reconciliation_remainder) CHECK --> GUARD : within tolerance, 2656..2661 (metrics.ledger_reconciliation_warning false) CHECK --> GUARD : unhappy, gap above half a percent or fifty tokens, 2656..2661 (warning true) PROV --> GUARD : 2680..2690 (provider_usage row; stage repo_escape_check) HASH --> ROLL : 2924..2936 (hash_after.ok; hashes_after exists) ROLL --> METW : 3020..3022 (metrics exists; metrics.stage_timings not null) METW --> LIVE : 3022..3023 (metrics exists; timing_current absent) LIVE --> MAN : 3025..3028 (timing_current absent; run_manifest cell entry exists)
Two states of the chapter’s own mechanism are not nodes of the model and are drawn in the timer’s diagram below instead: the open span and the closed span of the stage timer. The timer’s states are those a span passes through, and the transitions are the four methods.
stateDiagram-v2 [*] --> Open_span : stage() 147 or open() 187 pushes _Span 71; _mark_live 125 rewrites timing_current.json Open_span --> Line_written_ok : block ended or close() 205; _append 118; measured true, ok true Open_span --> Line_written_failed : exception inside the block, 164..166; ok false, exception re-raised Open_span --> Lost : close() never called (open without close, 194..197); no line [*] --> Reported_line : record() 229; measured false; started_at and ended_at as given or ended now Reported_line --> Nothing : seconds is None, 254..255; no line Line_written_ok --> Roll_up : read_timings 337, summarize 352; containers 333 excluded from accounted_seconds Line_written_failed --> Roll_up Reported_line --> Roll_up : counted in breakdown_seconds, never in accounted_seconds, 371..375 Roll_up --> [*] : metrics.stage_timings, run_cell 3020; rewritten by score_cell 1403 Line_written_ok --> [*] : clear_live 270 deletes timing_current.json at run_cell 3023, 1863, 1883, 1957 and score_cell 1399
6. The sequence of one pass’s record chain
The participants are the runner, the timer, the transcript readers of the timing module, the ledger builder, and the files. The route shown is the Fabric route on a Claude-shaped back end; the Anthropic route omits the four lifecycle lines, and the Codex route replaces the ledger with the provider usage line.
sequenceDiagram participant R as run_cell.run() 2464 participant T as StageTimer (timing_lib 88) participant TL as timing_lib readers participant LL as ledger_lib.build_ledger 420 participant X as cell files R->>T: open("pass", k) 2466 T->>X: timing_current.json (_mark_live 125) R->>T: open("agent_invocation", backend, containerized, prompt_chars) 2503 Note over R: the agent runs (chapter 06, section 4.10 c) R->>T: close(agent span, ok, returncode) 2547 T->>X: timings.jsonl line agent_invocation (_append 118) R->>T: record lease_wait, lease_admission, agent_process, lease_release 2566..2573 T->>X: four lines, measured false R->>TL: transcript_timing(stream) 2575 TL->>X: read pass k .stream.jsonl (_stream_events 433) TL-->>R: session, model_api, ttft, gaps, by_tool, requests R->>TL: write_turn_ledger(stream, turns path, test_command) 2589 TL->>X: transcripts/pass k .turns.jsonl (1424..1430) R->>TL: read_turn_ledger 2591, ledger_summary 2597, request_summary 2598 R->>T: record agent_model_api, agent_own_commands, agent_model_gaps with the span's start and end 2607..2634 T->>X: three lines, measured false, derived_from transcript R->>X: command_log.jsonl 2636 R->>T: stage("build_token_ledger") 2647 R->>LL: build_ledger(stream, k, injections, lap_manifest) 2649 LL->>X: read pass k .stream.jsonl (_read_stream 267) LL->>LL: _normalize_stream 276; step A 431; step B 441; classify_event 169; step C 526 LL-->>R: rows (injections, tool calls, messages, remainder) T->>X: timings.jsonl line build_token_ledger R->>X: token_ledger.jsonl append 2651..2653 R->>R: pass_spend 2654; _reported_usage 2656; warning on gap 2659..2661 R->>R: cumulative += pass_spend 2679 Note over R: guards, tests, patch, hidden checks (chapter 06) R->>T: close("pass") at the next iteration 2465 or at 2917 Note over R: finalisation R->>X: category_totals from token_ledger.jsonl 3013 (3057) R->>TL: read_timings and summarize 3020 (3032..3045) R->>X: metrics.json 3022 R->>T: clear_live 3023 T->>X: timing_current.json removed (270)
7. The records, with their writers and readers
Every file this chapter’s code writes, the line that writes it, and every program that reads it. metrics.json is chapter 06’s record; the rows here name only the fields this chapter’s mechanism fills. The campaign timing tables are chapter 10’s records; their rows here name only the writer this chapter reaches. Readers are named by file and line where the reading is one call and by file where it is spread. A column for the agent’s role is not needed yet: every record here is single agent only, and chapter 14 will add the role where a record would split.
| Record | Writer | Fields or content | Readers |
|---|---|---|---|
cells/<cell>/timings.jsonl (flow-model key stage) | timing_lib.StageTimer._append 118, through stage 147 to 185, close 205 to 227 and record 229 to 268; opened by run_cell.py 1806, score_cell.py 1277 and 1440, run_batch.py 461 | Per line: schema (5b.cell_timings.v001), run_id, cell_id, stage, phase (setup, pass, finalize, scoring, batch), pass, parent, depth, started_at, ended_at, seconds, ok, measured, attrs (free per stage: backend, containerized, prompt_chars, returncode, rounds, derived_from, time_to_first_token_seconds, turns, the request summary and ledger summary keys, by_tool, tool_calls, test_runs, test_seconds, sleep_seconds, resume_number) | timing_lib.read_timings 337, called by run_cell._timing_summary 3043, score_cell._refresh_stage_timings 1419, timing_lib.pass_timing_rows 1530, aggregate_timings.cell_rows 499 and collect 1082; server.py _cell_started_epoch 909 to 921 (the first line’s start stamp, because the metrics record is only written at the end) |
cells/<cell>/timing_current.json (timing_current) | timing_lib.StageTimer._mark_live 125 to 145, at every stage, open and close; deleted by clear_live 270 to 275 at run_cell.py 3023, 1863, 1883, 1957 and score_cell.py 1399 | schema, run_id, cell_id, updated_at, stage, phase, pass, started_at, stack (the names of every open activity, outermost first) | server.py _current_cell_detail 1482 (the live marker constant at 1479) and the current-stage block at 1958 to 1974; _cell_started_epoch 921 as a fallback. Nothing reads it after the cell ends |
cells/<cell>/transcripts/pass<k>.turns.jsonl (turn) | timing_lib.write_turn_ledger 1399 to 1431 (the write at 1424 to 1430), called by run_cell.py 2589 and backfill_turn_ledgers.backfill 147 | Per line, one model request: request (from 1), started_at, first_block_at, completed_at, prompt_tokens, cached_tokens, output_tokens, reasoning_tokens_estimated, sub_agent, blocks (thinking, text, tool_use counts), tool_calls (each with name, started_at, ended_at, seconds, result_bytes, is_error, and for a shell call command truncated to 120 characters, runs_tests, sleeps, or for a file tool file_path), wait_to_first_block_seconds, stream_seconds, round_trip_seconds (built at 866 to 1051) | timing_lib.read_turn_ledger 1434, called by run_cell.py 2591, timing_lib.pass_timing_rows 1559 and server.py timings_pass 3347; aggregate_metrics.compute_cell 539 to 541 through pass_timing_rows; aggregate_timings.collect 1098 through the same |
cells/<cell>/token_ledger.jsonl (ledger) | rows built by ledger_lib.build_ledger 420 to 539 and appended by run_cell.py 2651 to 2653; rewritten whole by regenerate_token_ledgers.rebuild_cell 91 to 107; removed before a fresh cell’s first pass at run_cell.py 2365 to 2366 | Per row: run_id, cell_id, pass_number, event_index (from 0 within a pass), event_type (harness_injection, tool_call, assistant_message), tool_name, target_path, category (one of the ten at ledger_lib.py 55 to 60), lap_artifact_id, input_tokens, output_tokens, cache_read_tokens, cache_creation_tokens, attribution_method (reported, estimated_chars4, estimated), notes (reconciliation_remainder pass <k>, unknown_tool_fallback:<name>, bash_fallback:<word>, or null) | run_cell.py 2409 to 2415 (resume carry-forward of the cumulative spend) and 3057 to 3062 (_category_totals_from_ledger); aggregate_metrics.load_cell 366 (_read_stream 197) and its gate at 1113 to 1120, which refuses a cell without one; regenerate_token_ledgers.rebuild_cell 95 to 97 (the old text, to detect change); gen_run_set_report.py 1583 (existence only); server.py 1569 (tokens spent so far, live view) and _ledger_pass_tokens 1758 to 1769 |
cells/<cell>/provider_usage.jsonl (provider_usage) | run_cell._append_jsonl 1396 at 2673 to 2677; removed before a fresh cell’s first pass at 2367 to 2368 | Per line, one Codex pass: run_id, cell_id, pass, provider (dgx_spark_fabric), model_alias, usage (the provider’s object, with total_tokens read at 2678 and 2418) | run_cell.py 2416 to 2421 (resume carry-forward). No campaign table reads it: the aggregator requires token_ledger.jsonl (1113) and regenerate_token_ledgers.py skips the back end (46 to 47) |
cells/<cell>/metrics.json fields filled by this mechanism (chapter 06, record table) | run_cell.py 3022: cumulative_spend, per_pass[].spend, per_pass[].cumulative, per_pass[].category_tokens, category_totals (3013), ledger_reconciliation_warning (2976), lease_wait_seconds (2969), stage_timings (3020); score_cell._refresh_stage_timings 1403 to 1424 rewrites stage_timings alone; regenerate_token_ledgers.rebuild_cell 79 to 89 rewrites the spend, cumulative, category and warning fields | stage_timings: schema, source (measured), accounted_seconds, harness_overhead_seconds, by_stage, by_phase, breakdown_seconds, per_pass (summarize 382 to 399) | aggregate_metrics.compute_cell 398 (the spend fields are recomputed from the rows and not read; lease_wait_seconds and seconds_to_success are copied at 632); update_monitoring.py 85 (counts the warnings); gen_cell_narrative.py 434 to 435 (prints the warning as an anomaly); server.py (the cell views) |
cells/<cell>/.token_ledger.jsonl.tmp, .metrics.json.tmp | regenerate_token_ledgers.rebuild_cell 103 to 107, then replaced over the real files | The rebuilt texts | none; they exist only between the write and the replace |
6. Metrics/cell_timings.csv, stage_timings.csv, pass_timing.csv, timing_summary.json (chapter 10, record table) | aggregate_timings.main 1155 to 1170, and server.py _timing_data 3272 to 3278 (the interface writes the same four files so an auditor who never opens a browser has them); the pass table’s columns are timing_lib.PASS_TIMING_COLUMNS 1466 to 1481 | built from this chapter’s records through read_timings 499, pass_timing_rows 1098 and summarize_campaign 868 | server.py 3280 to 3363 (the three timing endpoints); the B4 forensics |
6. Metrics/cell_factor_matrix.csv token and timing columns (chapter 10, record table) | aggregate_metrics.write_all 1152 to 1155, from compute_cell 398 | input_tokens, output_tokens, cache_creation_tokens, cache_read_tokens, tokens_to_success, doc_maintenance_tokens, lap_read_tokens, tokens_to_success_docs_subtracted, attribution_exactness, reconciliation_remainder_share, cache_hit_ratio, the thirteen cell timing columns from cell_timing_from_passes 1633 to 1681, input_tokens_new_context, tokens_to_success_new_context, sub_agent_requests, attempts_with_compaction | chapter 10 and chapter 12 |
The lineage from these files to the campaign tables and the interface is drawn below. Every arrow is a reader relationship from the table above.
flowchart LR subgraph cell [cells/<cell>/ written during the cell] TR[transcripts/pass k .stream.jsonl<br/>run_cell 319 and agent_backends 979] TI[timings.jsonl<br/>timing_lib._append 118] TC[timing_current.json<br/>timing_lib._mark_live 125] TL[transcripts/pass k .turns.jsonl<br/>timing_lib.write_turn_ledger 1424] LG[token_ledger.jsonl<br/>run_cell 2651 from ledger_lib.build_ledger 420] PU[provider_usage.jsonl<br/>run_cell 2673] M[metrics.json<br/>run_cell 3022; stage_timings rewritten by score_cell 1403] end TR --> TL TR --> LG TR -->|transcript_timing 445| TI TL -->|ledger_summary 1054 into attrs| TI LG -->|_category_totals_from_ledger 3057| M TI -->|summarize 352 at 3020| M RG[regenerate_token_ledgers.rebuild_cell 43] TR --> RG M --> RG RG -->|rewrites 103..107| LG RG -->|rewrites 103..107| M BF[backfill_turn_ledgers.backfill 112] TR --> BF BF -->|write_turn_ledger 147| TL AG[aggregate_metrics.compute_cell 398<br/>load_cell 366; pass_timing_rows 539] LG --> AG M --> AG TL --> AG TR --> AG AT[aggregate_timings.collect 1054<br/>cell_rows 499; pass_timing_rows 1098] TI --> AT TL --> AT TR --> AT M --> AT CFM[(6. Metrics/cell_factor_matrix.csv<br/>write_all 1152)] PCM[(pass_count_matrix.csv 1156)] TM[(time_matrix.csv 1158)] AG --> CFM AG --> PCM AG --> TM CT[(6. Metrics/cell_timings.csv, stage_timings.csv,<br/>pass_timing.csv, timing_summary.json<br/>aggregate_timings.main 1155..1170; server 3272..3278)] AT --> CT UM[update_monitoring.py 85<br/>ledger_warning_count] M --> UM SV[server.py: _cell_started_epoch 909, live detail 1482 and 1966,<br/>_ledger_pass_tokens 1758, _timing_data 3247, timings_pass 3330] TI --> SV TC --> SV LG --> SV TL --> SV CT --> SV SV --> UI[the results interface]
8. The loops and the waits
Nothing in this chapter waits on anything external, and no function retries. The loops are bounded by the size of the files they read. _normalize_stream and build_ledger each walk the transcript once (ledger_lib.py 310 to 328 and 447 to 522), and _reported_totals walks it once forwards and once backwards to the last result event (393 to 411). transcript_timing, request_records and turn_ledger each walk the transcript once (timing_lib.py 495 to 553, 650 to 748, 878 to 1030). summarize walks the timing rows once (365 to 383), and pass_timing_rows walks them once per pass number found in the file or in the transcript folder (1533 to 1539, 1546 to 1548). The two offline tools walk every run set, every cell folder and every transcript in sorted order (regenerate_token_ledgers.py 114 to 122 and 133 to 141; backfill_turn_ledgers.py 125 to 157). The only child process is the timing aggregator the backfill starts with --aggregate (201 to 204), which has no timeout because it runs after every ledger is already on disk and its failure changes nothing that was written (the comment at 198 to 200).
9. The guards and refusals
| Guard | Where | What it refuses or records |
|---|---|---|
| A timing write can never raise | timing_lib.py _append 118 to 123 (contextlib.suppress(Exception)), _mark_live 125 to 145, record 256 (suppresses TypeError and ValueError), clear_live 270 to 275 | A full disk, a read-only mount or an unserialisable value loses the line and nothing else |
| A failed stage still leaves its seconds | stage 164 to 185 | The line is written with ok false before the exception is re-raised unchanged |
| A reported figure with no value writes nothing | record 254 to 255 | seconds of None produces no line, so an absent report is never recorded as zero |
| A reported figure is never a stopwatch reading | record 266 (measured false); summarize 371 to 375; harness_overhead_seconds 1391 to 1394; pass_timing_rows 1578 to 1586 | Lines read out of the agent’s report are reported as a breakdown and never added to a total |
| A containing activity is never added to a total | CONTAINER_STAGES 333 to 334; summarize 374; harness_overhead_seconds 1393; pass_timing_rows 1584 | pass and the four driver views are drawn but not summed |
| A damaged file is read as far as it goes | read_timings 337 to 349; _stream_events 433 to 442; read_turn_ledger 1434 to 1461 | A missing file yields no rows; a half-written last line is skipped rather than raising |
| A turn ledger that cannot be written stops nothing | write_turn_ledger 1417 to 1430; run_cell.py 2588 | An unreadable or empty transcript writes no file and reports zero requests; the runner suppresses any exception around the call |
| An unknown tool never crashes the ledger | ledger_lib.classify_event 205 to 206; build_ledger 499 to 500 | Rule T1: the row is reasoning_out with method estimated and a note naming the tool |
| A shell command nothing matched is a code read with a note | _classify_bash 147; build_ledger 501 to 504 | Rule B5: bash_fallback:<word>, so the fallback is countable |
| A partial edit that will not parse is a code edit | _is_comment_only_edit 84 to 91 | A SyntaxError or ValueError from the stripper means the edit is not provably comment-only and is classified conservatively |
| A missing artefact manifest classifies everything as ordinary code | load_lap_manifest 224 to 225 | The STR and NON builds, which carry no manifest, get empty sets and no lap_read or doc_maintenance rows |
| The rows always sum to the reported usage | build_ledger step C 526 to 538 | A non-zero remainder in any field is written as one marked row rather than lost |
A synthetic per-message output is never labelled reported | _normalize_stream 365 to 376; build_ledger 486 to 489 | Output distributed from the terminal record is estimated_chars4 even for a single block |
| An error result’s zeros never erase usage | _reported_totals 403 to 410 | The per-field maximum of the assistant sum and the terminal usage is taken, never the terminal figure alone |
| A large unattributed share is flagged | run_cell.py 2656 to 2661 | ledger_reconciliation_warning when a field’s gap exceeds half a percent or fifty tokens |
| The offline rebuild refuses a ledger that does not reconcile | regenerate_token_ledgers.rebuild_cell 74 to 79 | ValueError naming the cell and pass; nothing is written |
| The offline rebuild touches only Claude-shaped cells | rebuild_cell 46 to 47; rebuild_run_set 118 to 119 | A Codex cell is skipped, because its record is not a transcript the algorithm can read |
| The offline rebuild writes atomically and only on change | rebuild_cell 95 to 109 | Temporary files replaced over the real ones; an unchanged cell is not rewritten |
| The backfill refuses to race a batch | backfill_turn_ledgers.main 180 to 186 (batch_is_running 59 to 77) | Exit 2 while active_batch.md says Status: running, unless --force or --dry-run |
| The backfill never rewrites a ledger the pass wrote | backfill 140 to 142 | An existing file is skipped unless --overwrite |
A roll-up of nothing is None, not zeros | run_cell._timing_summary 3033 to 3045; timing_lib.ledger_summary 1111 to 1113; cell_timing_from_passes 1647 to 1648; new_context_summary 1310 to 1311; new_context_tokens_total 1354 to 1363 | A cell without timing, a pass without a ledger, or an attempt without a figure reads as unmeasured and never as zero |
10. Every unhappy path, in four parts
Each row states what the step is supposed to do and why it works that way, what goes wrong with the trigger and the code path, what is written with the status and exit, and what it costs downstream for the driver, the scorer, the aggregator and the interface.
| Supposed to do, and why it works that way | What goes wrong: trigger and code path | What is written; status and exit | What it costs downstream |
|---|---|---|---|
| Append one timing line per finished activity, so the cell’s wall clock can be divided; wrapped so a diagnostic can never destroy a measured result | The disk is full, the mount is read-only, or a stage attribute will not serialise; timing_lib._append 118 to 123 or record 256 suppresses the exception | The line is lost; the cell continues; no status change | stage_timings in metrics.json is incomplete for that activity; aggregate_timings.cell_rows 499 reads the file as measured and the missing activity is absent from stage_timings.csv, not reconstructed; the interface’s timeline has a gap |
| Rewrite the live marker at every boundary so the operator’s live view can name the activity in progress | The same write failures; _mark_live 128 suppresses | The marker is stale; nothing else changes | The live view (server.py 1966 to 1974) names the wrong activity until the next successful rewrite; no record is affected |
| Remove the live marker when the cell finishes, so its survival marks a crash | The cell crashes or is stopped before run_cell.py 3023, or clear_live 270 fails | timing_current.json survives with the last activity named | The interface reads a running activity for a cell that is not running; the flow model’s witness absent(timing_current) on the edge cell.finalising.metrics_written->timing.live_marker_cleared fails, which is the intended signal |
Time a stretch of code with open and close where a with block would mean re-indenting load-bearing code | A caller returns early without close; timing_lib.py 194 to 197 state that the line is lost rather than noticed | No line for that activity; the span stays on the stack, so later lines carry it as parent and a wrong depth | The roll-up lacks the activity; the interface’s nesting is wrong for the rest of the cell. Every open in the runner is paired with a close on every path that was read (2096 to 2132, 2137 to 2158, 2204 to 2354, 2465 to 2466, 2503 to 2547, 2917); a new early exit added between an open and its close would reintroduce the loss |
| Split the agent’s turn into model time and command time from the transcript, and place the three figures inside the turn | The transcript is missing, empty or unreadable; _stream_events 433 to 442 returns no events, transcript_timing 491 to 492 returns nulls | record writes nothing for a None figure (254 to 255); the agent_model_api, agent_own_commands and agent_model_gaps lines are absent | pass_timing_rows leaves model_wait_seconds and tool_seconds null (1585 to 1589); the flow model’s edge ledger.turn_ledger_written->timing.agent_report_recorded has no witness and the analyser reports residue |
| Write the turn ledger beside the transcript so the request-by-request view is a read rather than a full parse | The transcript holds no request, or the write fails; write_turn_ledger 1417 to 1430 returns 0; run_cell.py 2588 suppresses | No pass<k>.turns.jsonl; the cell continues; the unhappy edge ledger.turn_ledger_written->timing.agent_report_recorded#2 is taken | pass_timing_rows rebuilds the ledger from the transcript and labels the row recomputed (1560 to 1566); server.py timings_pass does the same (3348 to 3353); backfill_turn_ledgers.py will write it later. Nothing is lost while the transcript exists |
| Build the token ledger from the transcript, deterministically, so it regenerates byte for byte | The transcript has a line that is not JSON; ledger_lib._read_stream 267 to 273 raises json.JSONDecodeError, and build_ledger at run_cell.py 2649 is not wrapped | The build_token_ledger timing line is written with ok false (stage 173 to 185); no ledger rows; the exception unwinds the runner; metrics.json is never written | The driver sees a non-zero return with no metrics record; the scorer, if run, records run_loop_did_not_finish (score_cell.py 1016); the aggregator refuses the cell for want of a ledger (1113); the cell has no disposition (chapter 06, section 10). No campaign cell has been seen to take this path; the transcript is written verbatim by the runner and is not edited |
| Attribute the reported usage exactly, so the rows sum to the transcript’s own totals | The rows do not sum to the reported figure in some field, which is the ordinary case when the model produced thinking tokens or the per-block counts were partial; build_ledger 528 to 537 | One more row, reasoning_out, with notes reconciliation_remainder pass <k> carrying the difference; the edge ledger.reconciliation_remainder->ledger.reconciliation_check#2 | reconciliation_remainder_share on the matrix row (aggregate_metrics.py 478 to 481) rises; attribution_exactness (475 to 477) falls; the aggregator’s reconcile_check 244 to 267 accepts the pass because the remainder row is present (257 to 258) |
| Warn when a large share of the pass could not be attributed | A field’s remainder exceeds half a percent of the reported figure or fifty tokens; run_cell.py 2659 to 2661 | ledger_reconciliation_warning true in metrics.json (2976); the cell continues, valid | update_monitoring.py 85 counts it; gen_cell_narrative.py 434 prints it as an anomaly; regenerate_token_ledgers.py clears it on a rebuild (89) whether or not the rebuilt remainder is still large. The specification’s section 8 says the discrepancy is also logged in the run set’s execution_log.md; the runner does not do that (the only occurrences of reconciliation in run_cell.py are the flag at 1402, 1819, 2661, 2976 to 2977 and 3104), so the promised log line does not exist |
| Read the reported usage the same way in the runner and in the aggregator, so the two checks agree | The two readers differ: ledger_lib._reported_totals 380 to 412 takes the per-field maximum of the normalised assistant sum and the terminal usage, while aggregate_metrics._reported_usage 270 to 279 takes the terminal usage alone when present and the raw assistant sum otherwise | Nothing is written by the difference itself | A pass whose normalised assistant sum exceeds its terminal usage (the v003 case the docstring at ledger_lib.py 12 to 19 describes in the other direction) would have rows summing to the larger figure and, if no remainder row was needed, could trip reconcile_check at 262 to 267 with a RuntimeError naming the cell, which aborts the aggregator’s run and therefore the closeout’s first stage (chapter 10). Whether any campaign cell meets this condition was not checked (section 14.3); the flow model’s open item 22 records the same gap |
| Record the Codex back end’s usage without passing it through the Claude algorithm | The provider reports no usage; run_cell.py 2678 would record zero, but chapter 06’s guard at 2757 to 2772 halts the cell first as usage_unavailable | provider_usage.jsonl gains a line with usage null before the halt (2673 to 2677); metrics.json invalid | The aggregator refuses the cell (no ledger, 1113); no Codex cell reaches a campaign table on any path, because the aggregator requires token_ledger.jsonl |
| Carry the earlier passes’ spend forward on a resume, because the budget is cumulative | A resumed Claude cell whose ledger file was removed or never written; run_cell.py 2410 to 2415 finds no file and carries zero | The new passes’ rows are appended to a fresh ledger; cumulative restarts | The stopping rule (2468) admits more spend than the budget; tokens_to_success from the aggregator is computed from the rows that exist and is silently low. Chapter 06 owns the resume; the row is here because the ledger is what makes the budget cumulative |
| Rebuild a ledger offline and prove it matches the transcript | The rebuilt rows do not sum to the reported usage; regenerate_token_ledgers.rebuild_cell 74 to 79 raises ValueError; or a pass has no per-pass metrics record (61 to 63); or a cell has no transcript (54 to 55) | Nothing written; the exception stops the whole invocation (main does not catch it), so later cells in the run set are not rebuilt either | The operator reads the message naming the cell and pass; the ledgers on disk are unchanged. --check mode has the same stops |
| Backfill turn ledgers without racing a batch | active_batch.md says a batch is running; backfill_turn_ledgers.main 180 to 186 | A message on standard error; exit 2; nothing written | The operator waits for the batch boundary or passes --force; the interface keeps rebuilding ledgers from transcripts on demand |
| Rebuild the campaign timing tables after a backfill | aggregate_timings.py exits non-zero; main 197 to 209 | The ledgers written stay written; a message on standard error; exit 1 | The four timing tables under 6. Metrics are stale until the aggregator is run by hand or the interface rebuilds them on its next request (server.py 3247 to 3278) |
| Recognise a shell command as a run of the product’s test suite | The command mentions a runner name in a comment or a printed string, or wraps the suite in a script the cell manifest did not register; _command_runs_tests 766 to 790 matches text only | runs_tests true or false on the tool call | test_runs, test_seconds and other_tool_seconds on the attempt row are off by that command; the docstring at 771 to 790 accepts the false positive because ruling it out would need re-running the agent’s commands |
11. The metrics this chapter produces and where each goes
The ledger rows are the aggregator’s only source for the token columns. compute_cell (aggregate_metrics.py 398 to 690) orders the rows, restricts them to passes 1 to P where P is the pass at which the scorer’s verdict became a success (404 to 407), and sums input_tokens, output_tokens, cache_creation_tokens and cache_read_tokens (414 to 418). tokens_to_success, the campaign’s primary metric (run_defaults.json key primary_metric), is the sum of input, output and cache-creation tokens over that scope, and None when the cell never succeeded (425); cache reads are excluded by the policy of token_ledger_spec.md section 9 and reported separately as cache_read_tokens and cache_hit_ratio (486). doc_maintenance_tokens, lap_read_tokens, tokens_to_success_docs_subtracted, search_to_edit_ratio and the navigation family come from the category column (420 to 446). attribution_exactness is the share of scoped tokens on rows whose method is reported (475 to 477), and reconciliation_remainder_share the share on remainder rows (478 to 481); both are the specification’s section 10 quantities. The runner’s own cumulative_spend, per_pass[].spend and category_totals are advisory: the aggregator does not read them, and regenerate_token_ledgers.py rewrites them from the rows so that they agree.
The timing record reaches two families of table. aggregate_metrics.compute_cell calls pass_timing_rows with the cell’s registered test command (539 to 541) and folds cell_timing_from_passes (timing_lib.py 1633 to 1681) into the matrix row (the thirteen columns requests, model_wait_seconds, tool_seconds, test_runs, test_seconds, failed_test_runs, failed_tool_calls, harness_seconds, machine_wait_seconds, reasoning_tokens_estimated, prompt_tokens_resent, largest_prompt_tokens, seconds_per_request_median, spread at 672), absent rather than zero when the cell has no attempt on record. The restated new-context figures are computed in the same call: input_tokens_new_context over every attempt and tokens_to_success_new_context over the success scope (564 to 566), through new_context_summary (1274 to 1326), which walks each request’s total prompt size (prompt_tokens plus cached_tokens where reported) and sums the growth from one request to the next, clipping a shrink to zero; and new_context_tokens_total (1329 to 1363), which returns None rather than a partial sum when any attempt in scope has no figure. The module comment at 1147 to 1271 records why these exist: the local model server behind the Fabric, vLLM 0.26.0, reports every request’s whole prompt as new input and never fills the cache field, so the existing input_tokens figure is biased against the local route by an unknown and probably large factor; the comment also lists four limits (compaction, retries, resumed passes, cache-creation tokens absent from the turn ledger). The harness README (1. Harness/README.md lines 353 to 359) states the same and records that the operator has not chosen which pair is the headline number.
aggregate_timings.py (chapter 10) reads the timing file line by line (cell_rows 499) and, for the cells that predate 2026-08-29, reconstructs the activities from artefact modification times with a source column saying which (docstring, 22 to 38); it writes cell_timings.csv, stage_timings.csv, pass_timing.csv (the columns at timing_lib.py 1466 to 1481) and timing_summary.json (1155 to 1170). It is not in the closeout chain; the results interface imports it and rebuilds the four files on its first request after a run set changes (server.py 3247 to 3278), which the flow model records as open item 16. stage_timings inside metrics.json is a convenience roll-up and no campaign table is built from it.
12. The tests that exercise the mechanism
The harness test suite under 5. Experiment/1. Harness/scripts/tests/ covers this chapter in six files. test_timing_lib.py (fifty-six tests, lines 46 to 999) covers the timer’s four methods, the never-raise rule (a write failure at 180, an unserialisable value at 191), the null timer, the container exclusion and the measured split of the roll-up, the damaged-file readers (299, 303), the transcript split and its recomputation (319, 335, 361), identifier-based grouping (403), the turn ledger’s fields and edge cases (539 to 775, including the unmatched tool call at 672, the 120-character truncation at 697, the missing transcript at 764), ledger_summary (795 to 832, including the instantly failing test run at 832) and the new-context arithmetic (877 to 999). test_build_ledger.py covers the reconciliation invariant in every field (68), the injection and category rows, the remainder row (121), the v002 and v003 stream shapes (179 to 231) and the upkeep edit reaching the ledger as doc_maintenance (256). test_classify_event.py covers every rule of the specification’s decision table by rule name (R1, R2, S1, E1, E2, E3, B1 to B5, A1, T1, lines 24 to 199). test_regenerate_token_ledgers.py covers the repair of a ledger and its cached metrics (11) and the skipping of a non-Claude cell (57). test_backfill_turn_ledgers.py covers the six behaviours of section 4.13 (54 to 133). test_aggregate_timings.py (lines 110 to 604) covers the reconstruction, the ranking and the four-way split of an attempt. test_integration_smoke.py asserts on a synthetic cell that the ledger reconciles (209 to 211), that the warning is false (206) and that stage_timings carries the setup and pass phases (224 to 237). The interface’s test_timing_views.py (10. User Interface/, lines 59 to 273) covers the request-by-request endpoint, the rebuilt ledger labelled recomputed (84), and the exclusion of an invalid cell’s seconds from the campaign figures.
Against the unhappy paths of section 10: the timing write failure, the unserialisable value, the missing transcript, the unwritten turn ledger, the remainder row, the offline rebuild’s refusal to reconcile, and the backfill’s refusal while a batch runs are each covered by a named test above. No test covers a transcript with a line that is not JSON reaching build_ledger through the runner, the stale live marker, an open without a close, the resume with a missing ledger, the difference between the two reported-usage readers, or the aggregator failing after a backfill.
13. The dated incidents that shaped the code
The timing module is dated 2026-08-29 in its first line, and the reason it exists is stated at lines 4 to 10: until then the harness recorded two wall clocks for a cell and could not tell thirty-five minutes of waiting on the model server from thirty-five minutes of re-running a slow test suite. The placement of reported figures on the cell’s clock is dated 2026-09-03 at line 243: until then every reported line was written with no start, and the results interface could list those figures but had nowhere to draw them. The turn ledger and the backfill are dated 2026-09-03 (backfill_turn_ledgers.py line 2, and the runner’s comment at 2577 to 2584: until then the records were computed, reduced to a handful of percentiles, and thrown away). The grouping of a reply by its message identifier rather than by adjacency records the bug it fixed: adjacency turned 45 replies into 58 in a run set 039 cell (timing_lib.py 643 to 647 and 822 to 829), and the same cell is the example at 612 to 615 of the first request sending five thousand tokens of context and the forty-fifth fifty-six thousand. The failed_test_runs figure carries the incident of the paid Anthropic route between run sets 016 and 066, where every shell command failed in under a second with a permission error because the container’s configuration folder was owned by root, so a cell read as three test runs that never ran (1085 to 1092; the aggregator’s shell_available column at aggregate_metrics.py 580 to 592 records the same). The reset of the clock after a reply with no tool call was a real bug caught in review (1019 to 1027). The new-context restatement is dated 2026-09-03 at 1148 and records that using prompt_tokens alone turned a hosted cell measured at 213,702 tokens into 2 restated tokens until cached_tokens was added back, checked on run sets 069 and 071 (1186 to 1193); the harness README at 353 to 359 gives the eight-cell check.
The ledger module carries two parser corrections in its docstring. Version v002 of 2026-08-28 (lines 20 to 33) was an operator-directed correction before registered local use: the command-line tool emits one assistant event per content block, each repeating the same usage snapshot, and v001 summed the repeats, inflating input by the blocks-per-request factor and recording zero output on the locally served stream; run set 035 seed 2 is the measured basis, where v001 read 8,240,883 input and 0 output against a terminal record of 3,802,470 and 22,403. Version v003 of 2026-09-02 (12 to 19) made a larger terminal figure the reconciliation authority because current streams carry small per-block output counts while the terminal result carries the complete total, and it made an error result’s explicit zeros unable to erase usage already present. The E2 rule for comment-only edits and the artefact identifier table are the specification’s (sections 6 and 7, cited at 85 and 68). The retirement of the single-model exclusion on 2026-09-03 (Decision Sheet item 51) is why a second model in the transcript is recorded and not judged at run_cell.py 2662 to 2667. The runner’s comment at 2641 to 2644 fixes that the Codex stream is never passed through the Claude algorithm.
14. The weakest claim, what was not checked, and the token line
The weakest claim
The weakest claim is the reader list for the results interface in section 7. server.py changed between the flow model’s commit and this one (51 lines in the diff), so its line numbers here were re-read at commit 6e2ee2c2b and will not match the flow model’s record entries (which say 901 and 3177 where this chapter says 909 and 3330); the currency check of the specification’s section 7 will catch the next drift, but the two documents disagree today. The second weakest is the statement that no campaign table reads provider_usage.jsonl, which rests on a literal search for the file name across the harness and interface folders; a reader that constructs the path from parts would be missed.
14.3 What was not checked
Whether any campaign cell has a pass where ledger_lib._reported_totals and aggregate_metrics._reported_usage disagree (section 10, the ninth row); finding out needs both readers run over every transcript, which is a results-integrity script under the B4 contract and not a design pass. Whether the execution_log.md line the specification’s section 8 promises for a reconciliation discrepancy was ever implemented anywhere other than the runner; only run_cell.py was searched. Whether score_cell.py line 1277 and run_batch.py line 461 are the only timers other programs open on a cell folder; capture_evidence.py was searched for StageTimer and found not to open one, but the interface was searched only for the readers. Whether the open and close pairs of the runner are matched on every early exit; the six pairs were read but not every return between them. The generation ledger specification was read for its section 6, which says the same build_ledger code path is used with the generation table swapped in; no script in the harness folder writes generation_ledger.jsonl today (a search of 1. Harness/scripts for the name returns nothing), so that specification describes a corpus-build procedure and not a mechanism this chapter can trace.
14.4 Token line
The session that wrote this chapter had consumed 236,446 tokens of context by the time writing began, measured as the difference of the remaining-token counter (15,000,000 at the start of the session, 14,763,554 at the last reading before the file was written), spent on the specification, the sample chapter, the flow model’s chapter 09 subgraph, the draft, the digest sections for the four files, the whole of ledger_lib.py and the two offline tools, timing_lib.py in four ranges, and the call sites in the runner, the scorer, the driver, the two aggregators and the interface. No harness script was run and no model tokens were spent by the harness.