The Two-Layer QA Report: Append-Only Sessions, Living Modules
Why end-of-day reconstructed QA reports fail, and how the two-layer format with machine-artifact-only numbers makes test evidence traceable.
TL;DR
A new testing rule fixes untraceable findings by requiring an append-only session report filled in live, with unique scenario IDs, machine evidence, and a test-data ledger. A living module file tracks contract changes in every commit. First session proved it: 31 scenarios, 27 PASS, all counts matched 13 database queries.
A test report rewritten at the end of the day repeats one defect: the numbers in its findings cannot be traced back. Evidence that should be a database query result or a network response turns into a transcript retyped from memory. Once the test environment changes and the test data is gone, re-running one scenario to verify a finding means guessing from the start. The owner decision now in force closes that gap with a fixed format: an append-only session report and a living module file. The format is codified in the testing README, so every manual and exploratory session follows the same structure without renegotiating it at the start.
Layer one: the session report, logged as it goes
The session report is created at the start of the session and filled in as each step executes, not reconstructed at the end. Its structure is mandatory: session metadata carries the date, the branch and HEAD commit, the environment with container status, the test account, and the method. The scenario matrix assigns a unique ID to every case, L-1 for a valid login or C-3 for a CRUD case, with the exact URL, numbered reproduction steps detailed enough for someone else to replay, expected versus actual, status, and evidence. One matrix row from a real session names its screenshot file as the attachment, so every PASS claim has an artifact to open, not an unattached word. Corrections never edit old lines; they arrive as new, marked lines.
The findings section uses sequential IDs such as R4-1 through R4-5, each carrying a severity, the code location down to the line number, machine evidence, impact, and decision options. The test-data ledger records every record or media item created, changed, and deleted, with database-query proof; cleanup must be proven, not claimed. The limitations section writes down what was not tested, so readers know exactly where the report stops. Every session is registered in the testing index, one row per session pointing at its report and screenshot folder, so no session disappears between reports.
Layer two: the module file that stays alive
The module file answers a different question than the session report: what the module's contract looks like now, and which design changes happened along the way. A session report is temporary and untouched once the session ends; the module file keeps being read and updated long after. Any slice that changes a page, form, menu, or endpoint must add a history row in the module file in the same commit. That layer is what keeps old findings interpretable months later: the module's latest status, its active findings, and the history of changes live in one place instead of scattering across a dozen session reports.
Two small rules tie it all together. Screenshots are named with the capture datetime in local time followed by the scenario ID, so the order and owner are readable from the file name. Every evidence reference in the report is a clickable relative link. And the most decisive rule of all: numbers may only come from machine artifacts, database queries, network captures, or JSON files, never from manual transcripts. That rule ends the old habit of recounting from the screen and copying the result into the report.
Proof from the first session
The first session under this format tested an admin information module: 31 scenarios, 27 PASS, 4 findings with IDs, and every UI count cross-checked against 13 fresh database queries with a full match. A separate retest walked the four baseline failures from before and none of them reproduced, with one new finding, a login flake, properly documented. The discipline spans domains: the OWASP Web Security Testing Guide demands structured evidence for every claim [2], and this format translates that demand into a daily routine. The cost is writing while working: the screenshot is named the same minute it is captured, the counting query runs before the browser tab closes. That is exactly what keeps the report believable once everyone's memory starts to fade, and the writing burden actually drops because the end-of-day reconstruction phase disappears.
Sources: [2] OWASP Web Security Testing Guide