Two-way document sync: the trustworthy tool refuses precisely
A pull returned different content than my push, with no error. The fix was not a smarter parser but a precise map of what the sync refuses to do.
TL;DR
My Markdown quietly broke after a pull because Plane sanitizes HTML and drops unsupported formatting without warning. Instead of forcing acceptance, I built a sync engine that refuses risky work, storing Markdown safely and using stripped text as the stable identity. It classifies every outcome into six explicit statuses and fails loudly rather than guessing, prioritizing data integrity over convenience.
The first time I ran a pull, the Markdown file in my repo came back different from what I had just pushed. No error message. No conflict detected. Just quietly missing formatting, replaced by blank space. My first guess was an encoding issue, or maybe a Git configuration mistake on my machine.
Back then I believed a good sync tool handles every format and rescues every bit of data. That guess died fast. After reading the server-side code, I learned the opposite: a sync tool you can trust is one that knows exactly when to refuse work, not one that guesses its way through broken markup.
The core problem is how Plane treats HTML. The server actively sanitizes description_html with nh3.clean against a strict tag and attribute allowlist, logging every removal [4]. Push an unrecognized rich-editor structure and the server silently drops it. The document lives in two places, Markdown in the repo and description_html in Plane, and nothing tells you when they started to diverge.
So the fix was not forcing the server to accept my format. It was building the refusal map before writing any merge logic. I store document bodies as escaped literal Markdown inside p and br elements, and rich-editor structures get refused outright instead of silently dropped.
An identity that cannot drift
The anchor is description_stripped: a plain-text field generated by the server, the stable identity plane [1]. Comparing DOM trees is fragile because the server rewrites them; comparing stripped text is not. Identity lives in visible text, with a one-line first-line marker, never in HTML comments that sanitization sweeps away.
The round trip itself stays tight through normalize_body canonicalization: extraction and parsing come back byte-exact. The same rule is restated in plane-init for sync parity, because skills cannot import across skill directories, so disciplined duplication plus a regression test holds the parity. When hashes disagree with no explainable base, the engine classifies it as an impossible-base refusal and stops; it never guesses which side is valid.
Six statuses and explicit decisions
I designed six possible statuses per entry. Not a plain log, but a decision map that forces explicitness about every anomaly: full success, conflict needing manual resolution, unsupported-format refusal, retrieve failure when a mirror item is missing, unexplained hash drift, and explicit abort.
When a meeting summary conflicts between local and server, the engine does not pick the newer version on its own. It produces a conflict summary that demands an explicit pull, push, or abort. The DECISIONS union merge only runs with exact-identity dedup and ISO sorting; it sounds rigid, but I trust documented failures more than automatic corrections that might be wrong.
One constraint has to be accepted with a straight face: a missing mirror item is a retrieve failure, not a recreate signal. The Community Edition has no archive over the API, so manufacturing data out of nothing is off the table. Failing loudly beats producing plausible garbage.
Proof in the field, not just theory
I hit a duplicate plain-text marker problem during render and parse. At first I blamed a sloppy regex. The real cause: extract_text and markdown_to_description were not truly lossless over the single-div served form the server hands back. Fixing that, then pinning it with a regression test, mattered more than any feature.
Every write to disk is atomic, so a half-written file cannot survive an interrupted run. The audit comment carries metadata only, never document content, so the hash comparison stays clean later.
This approach may feel rigid. Plenty of people prefer flexible tools that quietly fix things behind the scenes. I pick loud refusals instead: tell me clearly what you will not do, and never trade data integrity for convenience that only looks safe.
Building a two-way sync engine taught me one thing: trust is not built by how smooth the happy path is, but by how consistently the system refuses risky work. When you design yours, do not ask "how do I handle every case?" Ask "which cases must I refuse outright?" The answer decides how safe your data stays.