Chapter 5.5 follows chapter 5.4, Anatomy of one task package, and precedes chapter 5.6, Scoring of a cell, within part 5, the task library.
The hidden tier contains checks that the measured agent does not see. A probe scores the behaviour required by the ticket without naming a code symbol. The ticket and the visible verification file together define the work. A hidden check may retest a stated criterion with inputs the agent has not seen. It may not add a requirement absent from the ticket. This keeps the check focused on general behaviour rather than on the shape chosen by the implementation.
5.5.1 Governing rule of the maintenance-task suite
A probe observes behaviour through an available surface. It does not require a module path, function name, class name, route spelling, fixture identifier, or error shape. This rule matters most when a feature task asks the agent to create routes, fields, or functions. The agent chooses those names, so a check that expects the reference names can record correct work as a failure. The authoring rule therefore makes the check discover the implemented surface and assess its result.
The rule also protects the witness. A witness records what a check can establish about the work. It must show the requested behaviour rather than reward an implementation that happens to resemble the reference build. The authoring guidance applies the same rule to generated address forms: an address must be derived from the subject being checked. Deriving it from a neighbouring subject can return the neighbour’s content and turn a refusal into a false pass. The target subject anchors the address.
5.5.2 Shared library modules of the maintenance-task suite
The checks share a library in the _hidden_lib folder. It provides discovery, tolerance, browser access, data readers, ranking, evidence handling, and handoff parsing. Task-specific checks keep the assertions from their tickets, while the shared library supplies the means of finding and observing the work. The library does not decide whether a ticket’s behaviour is correct.
The library contains ten modules. source: 5. Experiment/4. Task Library/_hidden_lib/ (folder listing)
fiveb_probe.pysupplies core probe scoring.fiveb_schema.pydiscovers the application’s published schema.fiveb_leniency.pyapplies tolerance rules for value matching.fiveb_judge.pycoordinates the language-model judge.fiveb_browser.pydrives browser checks of rendered pages.fiveb_read_board.mjsreads board state.fiveb_read_rows.mjsextracts rows from table structures.fiveb_rank.pyorders results.fiveb_ladder.pyrecords evidence levels.fiveb_handoff.pyparses the agent’s closing note.
5.5.3 Evidence levels of the maintenance-task suite
The evidence ladder has four levels. source: 5. Experiment/4. Task Library/_hidden_lib/README.md line 51
The first level calls the application’s interface inside the scoring process. It is fast and repeatable. The second level reads the data from which the page is built. It can be unavailable for a single-page application when the returned page carries no work item data. The third level drives the rendered page in a browser and reads what a person can see. It can catch a value that reaches the data but never reaches the rendering. The fourth level asks a language model to assess what the browser saw.
The deterministic levels run for every cell so their results can be compared. The language-model level is held back until the deterministic levels have failed to confirm the behaviour. A held-back judge is recorded as held back, not failed. A deterministic pass is settled before the judge is consulted. A judge-only confirmation is marked separately so its cost and uncertainty remain visible.
Figure D-T4-1. The surfaces a hidden check may name and the surfaces it may not
%% figure D-T4-1 flowchart LR n1_1["_hidden_lib/README.md"] --> n1_2["PROBE_AUTHORING.md"]
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section T4. Purpose of the hidden tier
The governing rule connects ticket behaviour to observation without requiring an implementation name.
Figure D-T4-2. A ladder diagram showing the evidence levels from interface observation through page data and browser rendering to a language-model judge, with the judge held back until deterministic checks fail.
%% figure D-T4-2 flowchart LR n1["A ladder diagram showing the evidence levels from interface observation<br/>page data and browser rendering to a language-model judge<br/>the judge held back until deterministic checks fail"]
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section T4. Purpose of the hidden tier
The evidence ladder orders deterministic observation before the language-model judge.
5.5.4 Language-model judge of the maintenance-task suite
The hidden_judge.json file configures the last resort. It selects whether the judge runs and which model performs the assessment. The judge sees screenshots and page text from the running application. It does not see source code, tests, or commit messages, because those would invite it to argue from the implementation instead of observing the result.
The judge can confirm behaviour after deterministic checks fail to confirm it. It cannot overturn a deterministic pass. Its use costs money and is not repeatable, so the result records whether the confirmation came only from the judge. This separation lets later analysis include or exclude judge-only confirmations.
5.5.5 Governing-rule failure of the maintenance-task suite
A probe once resolved internal symbols by name, compared prose case-sensitively, and derived an expected document outline from one reference data shape. Agents produced different shapes that met the ticket, yet the checks recorded failures. The mechanism was dependence on reference names and formatting rather than on the requested behaviour. The evidence was the set of conforming implementations rejected by the probe. The cost was wasted agent work and incorrect failure records. Rewriting the probe to discover the surface and score its result removed that source of failure.
5.5.6 Address-anchoring failure of the maintenance-task suite
A generated address was derived from neighbouring headings instead of from the requested subject. When a section was tampered with, the check could read an untampered neighbour and present that content as the subject’s excerpt. The mechanism allowed a nearby address to answer for the target. The evidence was a refusal that became an apparent valid excerpt. The cost was a false pass that could have approved a broken check. The anchoring rule now restricts derivation to the target subject.
5.5.7 Chained-baseline failure of the maintenance-task suite
A discovery check compared a chained step with the tree handed to that step. That tree already contained surfaces created by the preceding step, so the comparison treated those surfaces as pre-existing and removed them from the difference. The mechanism used a step-local baseline where the chain required its original baseline. The evidence was an absent checklist behaviour reported even though the predecessor had produced it. The cost was failed cells and skipped later work. The baseline rule now uses the tree from which the whole chain began.
Figure D-T4-3. An incident diagram showing three check failures: symbol-based probing, neighbour-derived addresses, and a step-local chained baseline, each leading from a mechanism to misleading evidence and an incorrect cost.
%% figure D-T4-3 flowchart LR n1["An incident diagram showing three check failures: symbol-based probing<br/>neighbour-derived addresses<br/>a step-local chained baseline<br/>each leading from a mechanism to misleading evidence and an incorrect cost"]
source: operations/site-ia/A3-diagram-specification.md
source: operations/site-ia/A6-page-briefs.md section T4. Purpose of the hidden tier
The incidents show how a check can misread correct work, accept a broken check, or suppress later evaluation.