Evidence
The published, read-only evidence for the campaign, gathered behind one destination. Prioritized analyses ranks the registered claims against the observations in this published snapshot; the two- and three-dimensional visualizations reuse the same principal component implementation the local research interface uses.
Scroll horizontally to compare the prioritized analyses. Select a card to update the projection below it.
Selected analysis
The published sky uses the same three-dimensional projection, controls, and filters as the local run monitor. It has no running-cell telemetry or programming-attempts panel because this page is a read-only snapshot.
Execution replay
A recorded cell timeline, shown with its source and provenance.
Experimental coverage
Counts of valid measured cells for every build and task pair in this published snapshot. A surrounding box marks tasks that must run in order: each step starts from the tree produced by the step before it, so a failure can stop the rest of that sequence. Empty later steps may be “unreachable, predecessor not measured”.
Research record
The campaign in date order: one entry for every significant finding,
result, instrument defect, incident and ruling on record, newest first. It is written by
hand and read from one file,
0. Plan/Appendix P: Experiment Log Summary.md, which summarises the lab
notebook, Appendix G: Experiment Log.md. The notebook is append-only and
is the record of what happened and when; this summary is derived from it and is
corrected whenever a deeper look changes what the campaign believes. Nothing on this
page is computed from the run sets, so every claim here is a judgment somebody wrote
and signed with a date, and each entry names where its full account lives. Entries
marked landmark open with their prose showing; the rest open on their heading,
their figures and one control to read the whole thing.
Current evidence record
Six phases of the campaign
Each phase is a date range read from this same record. Selecting one filters the entries below without losing the search, kind, weight, topic or order controls; select it again, or use "every phase", to see the whole record.
How this page is kept true, and the terms it uses
Run set reports
One report for each run set. The first part is assembled here from the run set's own records, so every figure in it can be pointed at the file it came from: the run manifest, each cell's scoring summary and metrics record, the batch specification, and the execution log. The second part is the written analysis, which may live in any of three places, and this page reads all three: the four write-up files stored beside the programming attempts, part 2 of the full results report, and the research article. An analysis a language model drafted at the end of the run set is shown as a draft nobody has read until a person accepts it. Keeping the parts apart is the run set template's own rule, that a reader must always be able to tell a measured number from a written judgment.
Where the time goes
Every measured cell is a sequence of separate activities: staging a working copy of the build, running the staging gate, each turn the agent takes, each run of the visible test suite, each run of the hidden checks, and the scoring that follows. Each one is timed on its own, so a slow cell can be attributed rather than guessed at. Cells run since the harness was instrumented carry their own stopwatch readings; earlier ones are rebuilt from their command log, lease lifecycle, transcripts and the modification times of their scoring outputs, and every row says which of the two it is. The request-by-request breakdown of one attempt is a local-only view and does not appear here; the panels below that read campaign-level and per-cell figures are unaffected.