Skip to content

The QA Spec Was Done by 9 a.m., Then the Owner Said: Too Narrow

Adityo Guni Waluyo

The QA backfill spec was pushed at 9 a.m. and widened by the owner minutes later: lock current behavior, measure a baseline, retest one full vertical slice.

TL;DR

Initial plan focused only on validator unit tests, but coverage alone doesn't prove correctness and misses cross-field validation. Owner shifted scope to measure baseline first, then retest one full vertical slice from database to UI with contract checks. Findings now go to an owner decision list instead of silent fixes, keeping characterization tests honest.

Just past nine I pushed docs/superpowers/specs/2026-10-09-qa-backfill-unit-slice-design.md for KotaPortal. Neat document, one focus: close the data-validation holes at the unit level. Validation in KotaPortal genuinely was not tested that deep, so starting at the lowest validators and services felt like the right move.

A chat message landed a few minutes later. One instruction: do not stop at unit depth. If we were going to backfill QA, retest one whole vertical slice, from the database to the page. The spec that had just been born suddenly felt too narrow.

My first guess: adding validator unit tests would be enough

My guess that morning was simple. Most validation bugs come from rules nobody wrote down, so the answer must be adding a unit test for every validator. I had the list of validators to cover in my head, chase the coverage number, done.

The problem is that this assumption treats coverage as correctness. It is not. Line coverage only says a line was executed by a test, even when that test has no assertions at all [1]. A 2014 study quoted on the same page found that 58% of catastrophic failures started from trivial mistakes that statement coverage testing could have exposed [1]. That number reminded me: high coverage without assertions still lies.

One part I had skipped. Data validation is not only a syntax check. OWASP is explicit: validation covers syntax, semantics, and consistency between related fields, and the rule is to define what the app accepts and reject everything else [2]. If I only chase validator units, I never answer cross-field consistency or the API contract the frontend already depends on.

I also forgot one thing. KotaPortal's own baseline was not honest yet. The spec notes that the per-package percentages on record are misleading because most of the suite lives in the integration tier. Without re-measuring, any number is just a guess.

The direction that arrived minutes later

The owner's directive inverted the order of work. Not "add tests first, measure later" but "measure first, then claim". That became Phase 0 in the revised spec.

Phase 0 is a baseline table, not a narrative. For Go I had to run the toolchain's own instrumentation through go test -count=1 -cover -tags=integration ./internal/... -p 1 plus the -coverprofile option it ships with [4]. For the frontend, Vitest is the yardstick, with a choice of the v8 or istanbul providers that instrument code differently [3]. On top of that sit run_page.sh <modul> per module and a smoke runner to see whether a page actually opens. Only once that table exists am I allowed to talk about gaps.

The change makes sense because it enforces one discipline: measure with the same tooling before touching anything. Without it, I would only be moving an uncomfortable feeling into a number that looks good.

The spec I ended up writing has four moves

The final spec is no longer titled "unit-level". I split it into Context, Goals, Phase 0 baseline, per-slice checklist, and Owner Decision List. The four moves lock into each other.

First, the characterization lock. With legacy code the source code is the truth, even when it contains bugs [1]. The spec says it explicitly: "Cover the code with characterization tests." [1] Current behavior is therefore correct first until the owner decides otherwise. No silent behavior change smuggled in through a new test.

Second, the Phase 0 baseline above. There are no precise percentage targets in this article, because internal numbers have no public URL to link to, so I keep it qualitative: service and validator coverage stays thin, integration dominates. Every number that can be verified still points at open sources, not at private documents.

Third, retest one whole slice, not one layer. The flow is DB → API → OpenAPI contract → FE → page. In the middle sits npx @redocly/cli lint so the OpenAPI description is machine-valid [5], and npm run gen:api to keep the contract in sync. Then run_page.sh <modul> and the smoke runner check that a page does not merely render but is reachable through the right routing. Two extra criteria land here: routing/link integrity so navigation does not break halfway, and a quality gate for the admin dashboard — CRUD, search/filter, plus loading/empty/error states that unit tests never covered.

Fourth, findings do not become fixes on the spot. Findings go to the Owner Decision List. If a test cannot be relied on to block a merge or a release, fix it or delete it [6]. And since E2E is expensive to maintain, its share stays minimal [7]. GitLab's recommended order is clear too: start at the lowest level, Unit → Integration → System → E2E [6]. The spec follows that, but it does not stop at unit.

I will be blunt here. A coverage number without a slice retest is a comfort metric. This spec is what turns "we should add tests" into a plan that can be executed and audited: a measured baseline, a per-slice checklist that can be ticked off, and a contract gate that can be re-run with the same command tomorrow morning. I closed the spec by moving every ambiguity into the Owner Decision List instead of into the code.

Sources

[1] https://tdd.mooc.fi/4-legacy-code

[2] https://cheatsheetseries.owasp.org/cheatsheets/Input_Validation_Cheat_Sheet.html

[3] https://vitest.dev/guide/coverage

[4] https://pkg.go.dev/cmd/cover

[5] https://redocly.com/docs/cli/commands/lint

[6] https://docs.gitlab.com/development/testing_guide/testing_strategy

[7] https://www.martinfowler.com/articles/practical-test-pyramid.html

Related articles