A 192-Page Crawl, One 404, One Wrong Citation
An internal link-audit crawler in Playwright: static seeds, dynamic discovery, and a wrong citation in the commit message replaced by verified sources.
TL;DR
A Playwright-based crawler hit one dynamic 404: a slug containing an apostrophe got double-encoded, breaking the URL before it reached the router. The cited Next.js issue was a false reference, so MDN docs and issue #89879 became the real evidence. Final tally: 192 pages, zero hard-broken links, fix scheduled separately.
The test run printed a summary that looked perfect: 192 pages crawled, zero hard-broken links, zero redirects, and one last line reporting dynamic404 = 1 on the path /pariwisata/sarana-olahraga/k'mang-futsal. The apostrophe in that slug was the first hint. The page exists in the database; what breaks is the way its URL is constructed on the way to the router.
The old habit at this point is to open the commit message and follow its reference. The crawler's commit pointed at issue 58421 in the Next.js repository as the double-encoding reference. A check through the GitHub API showed that number belongs to a closed pull request that adds sorting-array polyfills. Nothing in it touches URL encoding. Followed blindly, that citation would have sent the fix after a problem that never existed.
The root cause that actually holds up
Two verifiable sources replace the broken reference. First, the MDN documentation for encodeURIComponent explains that the function escapes a larger set of characters, and that encoding an already-encoded string escapes the percent sign again, turning the code for an apostrophe into a nested code the server does not recognize [9]. Second, open Next.js issue #89879 documents internal request URLs being re-serialized incorrectly until query values split apart [8]. Together they explain how a slug with a special character lands on a 404 while its data sits in the CMS. The specific interaction between this application's API client encoding and Next 14 parameter handling stays a working hypothesis in the owner-decision ledger, not a verified fact.
The crawler pattern: static seeds, dynamic discovery
The crawler lives in a 58-line spec file. Fourteen public static routes form the seed list; everything else arrives through a BFS queue built from the href attribute of every link inside the application's origin. A goto call returns the main resource response, so the status and the final URL are readable immediately [4], and an in-page evaluation extracts every href from the open page [4]. Three filters keep the crawl on rails: a visited set prevents fetching the same page twice, hash fragments are skipped because they never reach the server, and URLs that do not start with the application origin are dropped.
const resp = await page.goto(url, { waitUntil: "domcontentloaded" });
if (!resp) { broken.push(`${url} (no response)`); continue; }
if (resp.status() === 404 || resp.status() >= 500) broken.push(`${url} -> ${resp.status()}`);
if (resp.url() !== url && resp.status() >= 300 && resp.status() < 400)
redirects.push(`${url} -> ${resp.url()}`);The results split into two classes with different consequences. hardBroken fails the test through an expect: a broken static route means the routing contract is violated and the build should go red. dynamic404 is reported only: dynamic routes depend on data that changes at any time, and reporting the miss is more honest than failing the whole suite over one stale content item. Every hard-broken page also leaves a screenshot through the evidence fixture.
Why a browser instead of static fetching
An HTTP-based link checker covers the server side and misses one layer: Next.js prefetches routes connected through its Link component when they enter the viewport, while dynamic routes skip that prefetch or only partially prefetch [6]. What a visitor experiences is not identical to the URL list in the sitemap. The sitemap protocol itself demands XML with entity escaping, a single host per file, and a mandatory loc for every URL [5]; the practices Google flags as most overlooked are size limits, file location, and which URLs belong in the file at all [7]. A browser crawler closes that gap because it walks pages the way a visitor does.
The final numbers for this session: 192 pages, zero hard-broken, one dynamic miss, zero redirects. One finding was enough to schedule the encoding fix outside this slice. The lesson that outlives the numbers: a reference in a work note is a hypothesis, not evidence, and a browser crawl that runs once in seconds is cheaper than one visitor finding the 404 first.
Sources: [4] Playwright documentation for the Page class [5] sitemaps.org protocol specification [6] official Next.js doc on linking and navigating [7] Google Search Central guide to building a sitemap [8] issue #89879 in the vercel/next.js repository [9] MDN reference for encodeURIComponent