Tests That Lock Design Intent, Not Screenshots
Design intent does not survive on documentation alone. A 129-line unit test pins the house classes, so drift breaks the build, not the quarterly audit.
TL;DR
When a fresh restyle kept slipping back to old borders and colors, thicker documentation didn’t stop the drift. So the team shipped a 129-line test that renders components to static markup and checks exact Tailwind classes. It now catches silent design changes in CI and enforces token-only styles before shipping.
Two weeks after shipping a data-component restyle, I opened the services page of the NusaLoka site just to admire it. Three PoCC mode cards, exactly as designed. Then I pulled a feature branch from a teammate, refreshed, and one card had quietly gone back to its old shape: thick border, plain white background, rounded-xl corners. I had not touched that file.
If you have shipped enough visual restyles, you know the arc. Week one, everything crisp. Week two, small deviations slip in with unrelated feature branches. Week three, someone "fixes" the spacing back to the old rhythm. Screenshots in the design doc age. Reviewer memory fades faster. The quarterly audit is the only thing that notices, and by then the drift has fingerprints all over the codebase.
My first guess was that this was a communication problem. So I did what documentation-minded people do: wrote a thicker design doc, with per-component class specifications and annotated screenshots. It changed nothing. Thicker documents have never once beaten thinner attention spans. The spec was right, and it was ignored, and both of those facts could be true forever.
On the latest restyle of seven data components to our house style, I tried the opposite investment. Once the design intent was fixed in the doc, I wrote a 129-line unit test file whose only job is to make the pattern impossible to drift past silently.
Checking strings, not pixels
The file, data-component-restyle.test.ts, needs no browser. Six components render through renderToStaticMarkup, which produces a non-interactive HTML string from a React tree [1]. That output cannot be hydrated, and for this purpose it does not need to be: I only want the markup.
The components use next/image, and the test does not care. One line replaces the module:
vi.mock("next/image", () => ({ default: () => null }));
import PoccSection from "./pocc";The ordering there is not a mistake and not luck. vi.mock calls are hoisted above every import [2], so the mock is in place before the component module loads, even though the mock line sits after the imports in source order. The first time I learned this it felt like cheating; now it feels like the whole point.
Assertions are string arithmetic. A tiny count() helper checks how many times an exact class string appears in the rendered markup, so the three PoCC cards must all carry the Pattern A tint rounded-2xl border border-primary/15 bg-card/40, and the four workforce cards must each lead with a size-8 text-accent icon:
it("pocc: mode cards use Pattern A tint with lead icons", () => {
const html = render(PoccSection, { content: pocc });
expect(count(html, "rounded-2xl border border-primary/15 bg-card/40 p-lg")).toBe(3);
expect(count(html, "size-8 text-accent")).toBe(3);
expect(html).not.toContain("rounded-xl");
});The negative lines matter more than the positive ones. not.toContain("bg-white") and not.toContain("rounded-xl") are the guards against the past. When someone pulls the old pattern back in from another branch, the test fails before anything ships, instead of the next audit noticing three months later.
Rules without enforcers are wishes
Our house rules are plain: tokens only, zero new hex, Pattern A tint for cards on tinted surfaces. A rule like that is only as strong as its enforcer, and discipline is a weak enforcer. Tests normally pin behavior. What makes this approach different is that the test pins a design contract instead. That only works because the classes are not arbitrary: Tailwind utility classes are generated from theme variables, the design tokens themselves [4], so the class string is effectively the public API of the design system. If it shifts, the design changed.
I know what the Testing Library principle says: the more your tests resemble the way your software is used, the more confidence they give you [3]. Asserting exact class strings is deliberately not that. It is a trade-off I take on purpose: the failure mode I am guarding is not broken behavior, it is silent drift. Behavior has its own tests. Design has this one.
Order still matters. The intent was decided first, in the design doc; the tests came after. Test-after works here because the test pins a decision that was already made, not a guess about implementation.
Results in CI: the suite passes 20/20 next to the build, and the diff grep shows zero new hex values and zero new font families. Those are the same numbers the TESTLOG records, except now a machine, not my good intentions, guarantees them.
So the question I ask has changed. Not "is this documented?" but "which test explodes if this rule is violated?" If the answer is none, the rule is not finished yet. Next on my list: one assertion to pin the token validator's pathspec coverage, and one to keep text-accent-text from ever coming back on card headings. Machines make better design reviewers because they never get impressed by a demo.
Sources: