The Markdown converter I trust is the one that refuses
A two-way repo-to-Plane converter that rejects anything outside a closed subset, line number included, instead of guessing.
TL;DR
The doc-sync converter rejects unsupported Markdown with an exact line number instead of silently truncating content. It only handles a strict subset like headings to h4, lists, tables and code blocks while frontmatter stays repo-side and images become pointers. This strict grammar and repo-side normalization makes round-trip conversions predictable and errors easy to fix.
I run the pull command of my doc-sync tool and get one line back instead of a page: line 57: raw HTML is not supported in markdown. No half-rendered result, no quietly missing section. The converter simply refuses, and it tells me the exact line where it gave up.
Years ago I would have called that a bug. A good Markdown parser is forgiving, or so I believed for a long time. It reads whatever it meets, the way most Markdown renderers do. That trust gets expensive: content that was silently truncated only shows up when you diff against the source, sometimes long after other systems have consumed it, and the one who tells you is often a reader, not the tool.
The module mdhtml.py in the plane-doc-sync skill was built on the opposite instinct. Its three functions, parse_md, render_blocks, and normalize_md, accept a closed subset only: ATX headings up to h4, paragraphs, lists nested by indentation, GitHub tables with escaped pipes and padded short rows, bold, italic, inline code, links, images on their own line, fenced code blocks with the language preserved, single-level blockquotes, and a horizontal rule. Standalone images render as a visible repo pointer, because Community Edition pages have no attachments. Everything else comes back as UnsupportedMarkdownError, a dedicated exception whose message always carries the line number, just like the one in my opening. YAML frontmatter never enters the conversion at all: it is repo-side metadata, split before parsing, never pushed, and reattached when pulling.
# mdhtml.py: code span contents are shielded first so a
# literal HTML example is not mistaken for markup
if _RAW_HTML_RE.search(_CODE_SPAN_RE.sub("\x00", line)):
raise UnsupportedMarkdownError(
"raw HTML is not supported in markdown", line_no,
)
The refusals come in flavors, and the test suite pins each one to its line: setext headings get told to use #, nested blockquotes bounce, an unclosed code fence is reported at the last line of the file, mixed list markers and separator-less tables refuse too. While migrating older documents, these errors doubled as a checklist: fix the line, run it again, done. Compare that with a debugging session that starts with "why is this page different" and hands you no hint at all.
A grammar that closes instead of opens
Markdown's ambiguity problem is old. The original syntax description was not unambiguous, implementations drifted apart for over a decade, and the nasty part is that nothing in Markdown counts as a syntax error, so the divergence often surfaces late [1]. Implementations still do not parse identical input into identical output [3]. CommonMark answered by standardizing one unambiguous grammar [2]. This converter takes the other road: it closes the grammar down to a small subset, so the parser has no room to guess. The contract is honest. What is not supported yet is rejected at the scene, with its line number, instead of being quietly improvised.
Canonicalization stays on the repo side
Plane is never a neutral pipe for the HTML we send: the payload gets sanitized, converted into the editor document format, and the whole body is replaced on every update; the docs even tell you to fetch the page after a write if you need to inspect what was actually stored [4]. Byte equality over the wire is impossible, and that is exactly why normalize_md lives on the repo side: comparisons run on the canonical form, not raw bytes. The guarantee is tested as a property in the suite, not demonstrated with one example: normalize_md(html_to_md(md_to_html(text))) must equal normalize_md(text).
My experience says a mirror tool that barks during parsing is the comfortable one to live with. A three-second error in a terminal beats auditing a tracker full of quietly truncated documents. If you build anything whose output another machine has to trust, make the parser brave enough to refuse. The funny part: since the subset got strict, I almost never see the error anymore, because disciplined documents never touch the fence.