Evidence

The published, read-only evidence for the campaign, gathered behind one destination. Prioritized analyses ranks the registered claims against the observations in this published snapshot; the two- and three-dimensional visualizations reuse the same principal component implementation the local research interface uses.

Execution replay

A recorded cell timeline, shown with its source and provenance.

Experimental coverage

Counts of valid measured cells for every build and task pair in this published snapshot. A surrounding box marks tasks that must run in order: each step starts from the tree produced by the step before it, so a failure can stop the rest of that sequence. Empty later steps may be “unreachable, predecessor not measured”.

Rates show passes or failures out of valid cells; zero-valid pairs remain gaps.

Research record

The campaign in date order: one entry for every significant finding, result, instrument defect, incident and ruling on record, newest first. It is written by hand and read from one file, 0. Plan/Appendix P: Experiment Log Summary.md, which summarises the lab notebook, Appendix G: Experiment Log.md. The notebook is append-only and is the record of what happened and when; this summary is derived from it and is corrected whenever a deeper look changes what the campaign believes. Nothing on this page is computed from the run sets, so every claim here is a judgment somebody wrote and signed with a date, and each entry names where its full account lives. Entries marked landmark open with their prose showing; the rest open on their heading, their figures and one control to read the whole thing.

Generated from the research record

Current evidence record

Reading current state…

Six phases of the campaign

Each phase is a date range read from this same record. Selecting one filters the entries below without losing the search, kind, weight, topic or order controls; select it again, or use "every phase", to see the whole record.

How this page is kept true, and the terms it uses

Run set reports

One report for each run set. The first part is assembled here from the run set's own records, so every figure in it can be pointed at the file it came from: the run manifest, each cell's scoring summary and metrics record, the batch specification, and the execution log. The second part is the written analysis, which may live in any of three places, and this page reads all three: the four write-up files stored beside the programming attempts, part 2 of the full results report, and the research article. An analysis a language model drafted at the end of the run set is shown as a draft nobody has read until a person accepts it. Keeping the parts apart is the run set template's own rule, that a reader must always be able to tell a measured number from a written judgment.

Where the time goes

Every measured cell is a sequence of separate activities: staging a working copy of the build, running the staging gate, each turn the agent takes, each run of the visible test suite, each run of the hidden checks, and the scoring that follows. Each one is timed on its own, so a slow cell can be attributed rather than guessed at. Cells run since the harness was instrumented carry their own stopwatch readings; earlier ones are rebuilt from their command log, lease lifecycle, transcripts and the modification times of their scoring outputs, and every row says which of the two it is. The request-by-request breakdown of one attempt is a local-only view and does not appear here; the panels below that read campaign-level and per-cell figures are unaffected.

What the campaign's hours were spent on

Inside one agent turn

the turn split into waiting for the model and running the agent's own commands

The campaign, by model and by build

on the left, how long one model request takes, by which model answered it; on the right, a typical attempt's seconds split the same five ways as the one-cell chart below, by which build the agent was working in

One model request, by model

One typical attempt, by build

First attempt against the retries that followed it

how much of the clock goes on passes after the first

One cell, activity by activity

pick a cell to see the shape of its life

Each attempt, five ways

One attempt, request by request

not part of the published copy; open the run set in the local research interface for this detail

Has it been getting slower?

median seconds per cell, by run set, split into the four parts of a cell's life

How the failure axes are found

The ordinary axes know nothing about which cells failed. These three do.

The first axis

It separates failing and passing cells while accounting for how much each group varies. The other two axes describe the remaining shape of the failing cells.

Measurements kept out of the fit

Measurements that are blank after failure, or that encode one execution route or build, are left out so the picture does not claim a distinction that the data cannot support.