Skip to content

Two-way converter: conservative on push, tolerant on pull

Adityo Guni Waluyo

A horizontal rule that vanished without a trace exposed a healthy two-way converter design: canonical push, tolerant pull that still refuses loudly.

TL;DR

My horizontal rule vanished on pull because the parser only handled <hr /> and silently ignored plain <hr>. The fix now throws a clear error for the bare tag instead of quietly dropping it. The tool now pushes conservative canonical HTML but only tolerates its known dialect on pull, backed by round-trip property tests.

I ran pull on my document sync tool and the output looked clean: no errors, no warnings. When I opened the page in Plane, the horizontal rule I had added manually in the editor was gone from the pulled markdown file. The element didn't turn into mangled text or broken formatting; it simply vanished from the file.

My first guess was the Plane editor stripping the element on save. I even suspected some janitor script running server-side. I spent a fair amount of time watching logs; everything looked normal. My guesses missed by a mile, and the last suspect I would have named was the pull parser of my own round-trip converter.

That parser is a subclass of Python's built-in HTMLParser, and here comes the surprise. Python routes <hr> and <hr /> through different callbacks: the XHTML-style form fires handle_startendtag [1], while the plain start tag lands in handle_starttag. The old code only tolerated the self-closing form. So the horizontal rule, written in the most common spelling, got swallowed silently: no exception, no log.

The conservative side and the tolerant side

The fix actually tightened things, not loosened them. The bare form now throws UnsupportedMarkdownError naming the HTML line number, and the parser still accepts only the self-closing form of that void element [2]. What changed is this: what used to disappear without a sound now refuses loudly. The rule is deliberate: pull never guesses.

The fix opened my eyes to a healthier two-way design overall. The tool mirrors markdown files from a .docs folder into Plane Pages and back into the repo. A good round-trip converter turns out to be asymmetric, and the asymmetry is exactly the lesson.

The push side acts as a conservative serializer. The inline_to_html function stashes code spans and links first, wrapping them in NUL characters, and only then escapes the remaining text. As a result, ampersands and angle brackets inside code spans or links get escaped exactly once. Plain text only goes through escaping for ampersands and angle brackets with quote=False, so quotation marks stay literal inside element text; only href attribute values escape quotes.

def inline_to_html(text):
    text = stash_code_spans(text)   # \x00c0\x00, etc
    text = stash_links(text)        # \x00l0\x00, etc
    text = escape(text, quote=False)
    return unstash(bold(italic(text)))

The emission is pinned too. Code fences carry a data-lang attribute, table cells serialize as escaped plain text, and images become a human-readable placeholder paragraph, because Plane Community Edition has no attachments. Pull simply maps that placeholder paragraph back to image syntax.

The pull side carries itself differently. It is tolerant, but only inside the dialect it knows. A div tag is treated as a transparent container, br inside a paragraph reads as a line break, and nested inline marks and nested lists un-nest back into markdown. That same dialect is what the push side produces, plus the tidying the built-in editor applies. Anything outside the dialect gets refused with an exception naming the HTML line number.

One more bug got fixed in the same commit, and this one is my favorite. The image placeholder mapping used to fire for any paragraph whose first line matched the placeholder text. The second and following lines of such a paragraph were thrown away silently. Now the mapping only applies when the paragraph is exactly one text run; anything longer round-trips verbatim.

Postel, split in two

Both cases connect back to Postel's law, the one everybody quotes from RFC 791: be conservative in what you send, liberal in what you accept [3]. The principle isn't wrong; the danger is the free interpretation, the kind that reads liberal as silently discarding whatever you don't understand. Silent tolerance only gets caught weeks later, usually when someone notices content shrinking.

This converter splits the principle by side. Push output stays canonical and conservative. The liberal attitude on pull only applies inside the known dialect; outside it, a loud refusal with a line number. So if the Plane editor someday changes how it tidies HTML, what shows up is an error naming the line, before the content itself changes.

The round-trip safety net

The safety net isn't ordinary unit testing either. There is a round-trip invariant tested as a property: markdown converts to HTML, converts back to markdown, and must normalize to the original text. The converter also runs over a real corpus of documents, the classic property-testing pattern of corpus plus oracle, exactly like the formatter-over-corpus example discussed by Hypothesis [4].

Two small bugs in one fix commit ended up producing a bigger design decision: tolerance doesn't mean accepting everything. If a two-way converter you rely on has ever "eaten" content without a trace, check which callback actually receives that tag first. Often the HTML is fine; the assumption about parser callbacks is what's wrong.

Sources

Related articles