All Unit Tests Green, Yet the Converter Still Broke
Three pull-side bugs slipped past unit tests that were green on each side; only the round-trip property test caught the unsynchronized two-sided contract.
TL;DR
A round-trip property test over real docs failed despite green unit tests, showing the push and pull sides followed different contracts. Fixes addressed nested bold links via a wrap stack, a regex stash leak fixed by looping, and empty quote markers needing canonicalization. The lesson is unit tests only prove local specs, so you need cross-boundary property tests.
The Wrong Guess, and What the Contract Really Was
That night I ran the property test that checks the converter's round trip over the real docs corpus. The expectation was simple: turn markdown into HTML, then back into markdown, and the result should be equal after normalization. The test failed hard. The terminal printed a series of diffs showing the format changed in a dozen files. Strangely, every unit test on both the push and the pull side was green.
The corpus was nothing exotic: a dozen real documentation files that get edited through the Plane editor every week, so this was no laboratory case. My first guess: one parsing function must be broken. I traced the logs line by line, hoping to find one if-else that mishandled an edge case. The guess was wrong by a mile. The problem was not one broken function; the two sides of the converter had pinned two different contracts. Each side was correct according to its own tests, but not a single test checked the contract between them.
The tool I built mirrors markdown files to a self-hosted wiki and pulls them back. Its core invariant: the markdown before and after must be equal after normalization. The first bug showed up in nested inline elements. The push side emits bold text wrapping a link. The pull side, an HTML parser built on html.parser [2], rejected it flat out with a nested inline elements are not supported message.
The pull-side test had pinned that rejection, while the push-side test had pinned the bold emission. Both passed, but combined, the round trip broke. The fix: the pull side now keeps a wrap stack of tag, collected text, and href. When the innermost tag closes, its markdown token composes into the parent's parts. I deleted the old pull-side test outright, because it pinned the wrong contract. When two tests disagree, one of them is pinning the wrong spec.
The Regex Trap and Escaping Rules
The second bug was more annoying. When a code span sat inside a link label, an internal token leaked into the final HTML output. I first blamed the parser logic, but the cause was in how re.sub works. Per the official docs, it only replaces the leftmost non-overlapping occurrences [1]. It never rescans its own replacements.
html = "link with <stash_1> code"
# re.sub only replaces leftmost non-overlapping
html = re.sub(r'<stash_(\d+)>', restore_code, html)
# if restore_code reveals another stash token, it survives
# fix: loop until stable
while "<stash_" in html:
html = re.sub(r'<stash_(\d+)>', restore_code, html)
A stash token revealed in the first pass survived, because the regex had already moved past it. The fix: loop the substitution until no stash token remains. This escaping context matters: CommonMark allows backslash-escaping any ASCII punctuation, but explicitly says escapes do not work inside code spans [3]. That is why the push side stashes code spans as tokens before escaping everything else.
Specs That Never Talk to Each Other
The last problem: the pull side kept quote lines containing only the greater-than character, while the push side drops them. The output diverged from the original markdown and made the diff dirty on every document update. Read the CommonMark spec and a block quote marker is just that character plus an optional space [3]. A line holding only that character is the quote's version of a blank line. The fix was to canonicalize it away, exactly like blank lines between blocks. One line in the parser, but the effect shows up in every file that has an empty quote.
Of the three bugs, the most expensive lesson for me: green unit tests only prove each part matches its own spec. They say nothing about whether the specs agree. It is easy to feel safe because the coverage report is green. But a green test is only valid inside its own module's bubble. Only a property test that crosses the component boundary, like a round trip over real files, can catch a two-sided contract drifting between two modules written in the same week.