Skip to content

Page identity without a hidden marker

Adityo Guni Waluyo

A v3.1 design moves page identity from a hidden HTML marker to a configured page_id, adds retrieve-after-write hash comparison, and keeps unknown forms fail-loud.

TL;DR

v3 kept page identity in an HTML comment in the body, so a single rewrite without it made verification block the page. Commit 67f50e0 specs v3-1 moving identity to config with strict naming and verifying writes via read-back hash comparison that fails loudly. It is design only, does not identify the stripping cause, and keeps tolerance tight to prevent drift.

The commit message I was reading says the v3 identity scheme broke on its first real migration push. Page identity was a hidden HTML comment marker inside the body, and one write came back without it, so verify classified the page as blocked. My first instinct was to blame the sanitization pass: Plane documents that it sanitizes HTML, converts it to the editor document format, and replaces the complete page body on update [1]. That made an aggressive server-side rewrite the easy explanation.

Identity Belongs Outside the Mutable Payload

Designing around a probabilistic guess is the wrong trade, so commit 67f50e0 moves identity out of the content and into configuration: a page_id in the config plus a strict naming convention. Config and version shape stay unchanged, so the migration path stays small.

The v3.1 specification recorded in that commit is a design, not a shipped implementation and not a production rollout. I want to be explicit about that, because the interesting part is the reasoning, not a deployment story. Probes in the same commit showed storage is otherwise verbatim and deterministic, which is exactly the condition under which a stable configured ID beats a marker embedded in the payload.

Retrieve After Write, Then Compare

The specification anchors verification on a read-back: fetch the page right after writing, compare normalized hashes, and treat a mismatch as a failure rather than something to repair. Python's HTMLParser gives the handler callbacks for start tags, end tags, text, and comments [2], which is the shape a normalizer needs when it walks retrieved HTML. Note the parser's own limit: it does not check that end tags match start tags, so normalization has to make its own decisions rather than assume well-formed input.

Tolerance is deliberately narrow. Only editor transformations that were actually observed and captured as fixtures are allowed through the comparison. Anything unrecognized fails loudly, because a silent auto-repair of unknown output is how a page drifts into a state nobody wrote.

The Part That Is Still Unknown

The commit does not identify the trigger for the rare server-side canonicalization that stripped the marker in the first place. I am not going to invent one. What the design does buy is containment: a configured ID survives renames, the read-back catches divergence immediately, and the fail-loud rule means the unexplained case shows up as a loud failure instead of an accepted-but-wrong page.

That is a narrower claim than a fix, and it is the claim the evidence supports.

Sources

Related articles