Skip to content

Unit Testing Whose Contents Are Not Just Unit Tests

Adityo Guni Waluyo

One script per test tier: vet, integration against real MySQL, a smoke gate, up to Playwright E2E.

TL;DR

Scattered manual tests were often skipped, so they were consolidated into four strict scripts for backend, frontend, E2E and smoke. Backend handles Docker and runs integration tests serially on real MySQL, while smoke validates deployment endpoints and E2E wraps Playwright with retries. README now defines what to run when, so a passing backend script signals the shared team contract.

# Unit Testing Whose Contents Are Not Just Unit Tests

It used to be that every time I wanted to push code, I typed a long sequence: go test -tags=integration -p 1 ./..., spin up Docker first for db-test, vitest run, sometimes Playwright with long flags. The order drifted depending on how much of a hurry I was in. And when chasing a deadline, the step that got skipped was usually the most important one.

My first guess was that the solution was better documentation. It turns out documentation is just decoration if execution stays manual. There was also a small irony in the folder name: unit_testing, even though its contents were not just unit tests. Four tiers at once live in there: unit, integration, E2E, and smoke.

This commit poured all of it into four scripts: one for backend, one for frontend, one for E2E, one for smoke. All of them are headed by set -euo pipefail [1]. The moment a single pipeline returns a non-zero status, the shell stops right there.

No more stories of a script continuing to run while an earlier step already failed.

Integration that talks to real MySQL

The part most often forgotten when typing commands by hand is the prerequisite. In run_backend.sh, Docker is checked before integration runs:

if docker info >/dev/null 2>&1; then
  docker compose -f docker-compose.test.yml up -d db-test
else
  log "Docker not running - integration tests will skip"
fi

go test -tags=integration -p 1 ./...

docker compose up -d is idempotent: services that are already running are not started again, and if the config changed it recreates them [5]. So this call is safe to repeat without fearing duplicate containers.

The -p 1 flag is not decoration either. All integration packages share one real MySQL 5.7 test database (SQLite is strictly banned), and -p controls how many test binaries may run in parallel, defaulting to GOMAXPROCS [3]. -parallel is a different flag: it only applies inside a single test binary [4]. Mixing up these two makes people insist they already changed the setting while turning the wrong screw.

Smoke as the deploy gate

run_smoke.sh answers the question "deployed, now what". It checks three URLs: the health endpoint, the root page, and the API root. The last one must return valid JSON, not just a 200:

curl -sf -m 10 "$BASE/health" >/dev/null || fail "/health not 200"
body=$(curl -sf -m 10 "$BASE/api/v1/") || fail "/api/v1/ not 200"
printf '%s' "$body" | python3 -c "import json,sys; json.load(sys.stdin)"   || fail "/api/v1/ is not valid JSON"

One failure means exit 1 and the deploy flow stops there. The default points at the local server; pass a staging URL as the first argument and it becomes a staging check. Smoke testing, which used to mean "manually poke it in a browser", is now a gate you can hook at the end of the deploy flow.

For E2E, run_e2e.sh is thin: npx playwright test "$@". That is the point. Every Playwright flag still works, including running a single spec file. What I tuned in the config is the retry: Playwright does not retry failing tests by default [2], and trace is set to on-first-retry so trace files are only recorded on the second attempt [6]. Debugging keeps its data; storage does not turn into a pile.

README guarded like a contract

The most valuable part turns out not to be the scripts but the when-to-run-what table in the README. Every save: backend in unit mode only. Before a commit: add the frontend lint. CI or a tag push: the full suite. After deploy: smoke. Before a release: E2E. So when someone asks "which tests do I run before merge", the answer is no longer up for debate.

I close with the same standard the README uses: if run_backend.sh is green, the code passed the same contract everyone else runs against. Not just green on my laptop.

Sources

Related articles