Baseline QA Before Fixes: Red Doesn't Mean Broken Code
The first QA baseline came back all red: a stale database volume and a healthcheck calling a missing wget, not a broken codebase.
TL;DR
Initial integration failures came from environment drift, not broken code, with a stale database volume and a missing wget in the healthcheck. Cleaning volumes with docker compose down -v fixed the setup and let all 28 API packages pass. Coverage gaps and E2E failures were recorded without patching to keep the baseline trustworthy for prioritized follow-ups.
I ran the first QA baseline from the work plan in docs/superpowers/plans/2026-10-09-qa-backfill-fase0-routing-admin.md, and the very first step immediately backfired: the terminal screen filled with red indicators, and all integration tests failed simultaneously. My first reflex was obvious—the latest code must be broken. I opened the latest commit diff, looking for syntax or logic errors that had just changed. There were none.
The problem lived in two places the code itself could not reveal. The test database held schema_migrations revisions newer than the code currently checked out: the database volume from a previous test session was still alive, while the running code had rolled back to an older revision. The migrator reads migrations and applies them sequentially; it does not guess what a person wants, and if the situation is ambiguous, it simply fails [3]. The second signal came from the Docker Compose healthcheck, which kept showing unhealthy even though the application was responding with HTTP 200: the healthcheck script was calling a wget binary that was not available in the development image.
At that moment, I did not immediately trust the red results. A baseline functions to record the reality of the system as it is, and recording means measuring something I do not alter midway. I only fixed the environment; any findings touching the code or test scripts went into a decision list so they would not get lost among quick fixes.
Environment drift is not a code bug
The standard teardown command in Docker Compose only removes containers and networks, while named volumes are preserved by default [1]. Volumes are designed as persistent data storage for the container machine: if none exists, compose creates it, and if deleted manually outside of compose, the next session creates it again [2]. That leftover volume from a previous session is what brought the old schema into the test.
The first command only cleans the environment: docker compose down -v. The -v flag removes the named volumes declared in the Compose file along with anonymous container volumes [1], so the old schema gets swept out with it.
Only after that did I run the integration tests exactly as step F0-1 in the work plan:
cd api && go test -count=1 -p 1 -cover -tags=integration ./internal/...
The result was clean: 28 packages passed, 0 failed, with per-package coverage printed in the output. If a coverage profile file is needed, Go provides it via go test -coverprofile and the built-in cover utility that reads it [4].
Drift like this is not born from a single wrong command, but from decisions that made sense at different times: keeping the volume so testing does not rebuild the database every session, then running code that has rolled back without cleaning the old volume. No component was broken, and that is why there was no error pointing to the component itself.
The numbers dictated the next work order. The bottom five packages sat at 0.0%, 20.4%, 30.7%, 31.6%, and 33.4%, while 18 packages remained below the target set by the team. That became my work list, organized by numbers, not by feelings about which part felt fragile.
Other layers are recorded, not fixed
The frontend layer was green from the start: 10 files and 69 vitest tests passed, lint showed 0 errors with 3 warnings, and the build produced 42 pages. Vitest itself calculates coverage via v8 or istanbul instrumentation [5], so the numbers were immediately usable. Smoke tests on three main routes also passed.
I recorded the E2E baseline exactly as it was: 45 pass, 3 fail, 1 flaky, 4 not-run. I did not fix a single one of those three failures during the baseline session. Changing a system while it is being measured makes the measurement numbers untrustworthy, so the classification (product bug or test bug) is deferred until that list of findings is fully reviewed.
One decision list after baseline
All layer results went into a single planning document, complete with checklists F0-1 through F0-6 checked off along with their status, not just PASS/FAIL. The next work order read directly from there: packages with the lowest coverage get patched first, the healthcheck is examined to see which binary is actually in the image, and the three E2E failures wait for the project owner's decision.
The two red signals I found were actually the cheapest to fix: one command for a clean environment and one line of healthcheck configuration. The rest required decisions, and the baseline kept those decisions firmly on the findings list, not in the hands of anyone who panicked at the sight of a red screen.
Sources
[1] Docker docs: CLI compose down
[2] Docker docs: Compose file volumes
[3] golang-migrate README
[4] Go docs: cmd/cover
[5] Vitest docs: coverage