Beat 011: the replay that found an analysis meaning two different things
Nothing persists between visits. The manifest is the only thing that leaves the browser and the only thing that comes back: a seed, some derived seeds, a clock, a configuration digest, an issue instant, an ordered list of the reader's edits, and now the build's own commit and the domain. No fields, no scores, no observations — those are derived, and replay is re-computation rather than the restoration of a snapshot.
That distinction is the reason the byte-identity test means anything. A snapshot compared with itself proves nothing.
The defect the replay found
The test that matters is the one a reader would perform: export from one browser context, import into another, and compare. It compares a digest of the run's fields and its analysed field, which the page shows so a reader can do the same by looking at two tabs.
It did not match. The seed matched, the step count matched, the instant matched. The analysis did not.
The shell built its analysis from the state as it stood when the run was constructed. For a fresh run that is the initial state, which is what the analysis instant was declared to be. For a replayed run it is the state after the replay has advanced it to the step the manifest records — so the same run, rebuilt from its own manifest, produced a different analysis from the one it had exported.
Neither run was lying about anything it said. The analysis panel simply meant something different in each of them.
The fix is one line: take the background before anything advances the state. The interesting part is that nothing else in the suite could have found this. The headless replay test compares run state and manifests, and both agreed. No single-visit test can tell an analysis-at-step-0 from an analysis-at-step-72, because a single visit only ever makes one of them. It took two visits and a comparison, which is exactly what AT-04 asks for and exactly the kind of acceptance test that looks redundant until it isn't.
And a gate that had been allowed to lie
Gate G-05 runs in a real browser. Playwright, outside CI, reuses an already-running preview
server by default — and a server left over from an earlier command serves the dist/ that
existed when it started. So the gate could ask its question of a build that is not the tree.
It did, once, during this beat: G-05 failed against a stale build and passed a minute later against a fresh one, with nothing changed in between. That is the worst kind of check — one whose result is not about the code.
The gate now sets an environment variable and the Playwright configuration refuses to reuse a server when it is set. A check that can pass for reasons nothing in the tree explains is worth nothing is the same sentence this project already had about checks that have never been seen to fail.
What the manifest is now
A strict schema, which is what makes "the manifest contains no field data" true rather than merely intended: a manifest carrying a field, a score or an observation is rejected for carrying a key the schema does not know, not because somebody remembered to look for those three words.
There is one definition of that schema — the zod object in the code — and a committed JSON Schema generated from it for anybody who wants to validate a manifest without this code, with a test that regenerates and compares. It is the same arrangement the data artefacts have under gate G-01, for the same reason: a copy nobody regenerates is a document about what the code used to do.
Four refusals and one warning
The checks happen before anything is provisioned, so a refused import leaves the run you have alone:
- the schema, naming the field at fault;
- the format version, checked before the shape, so a manifest from another version is told what it is rather than told about a field that moved;
- the configuration digest — refused, because the declared values differ and a run made against other values is a different run;
- the domain, naming the one the manifest asked for and the ones this build has.
And one that is deliberately not a refusal. A code-version difference is a warning printed beside the run: a reader holding a manifest from last month is better served by a replay that says identity is only promised for the same code than by a door.
Nothing persists, and now it is measured
NFR-02 says nothing persists between visits — no storage, no cookie, no run in the URL. That was
a claim about intent until this beat. A test now replaces localStorage.setItem,
sessionStorage.setItem, document.cookie and indexedDB.open before the page loads, drives a
whole visit — integrate, build the row, draw a new run — and asserts that nothing was written.
Where it stands
241 headless tests, 34 shell tests, seven gates. One beat remains: 012, adaptive sampling, which is deferred behind a trigger the SRD wrote and which this project has not met.