One Findings List for One Owner Decision
Four QA reports became one master table: 18 findings awaiting a decision, 4 not reproduced. Who decides matters more than ticket age.
TL;DR
Four QA reports consolidated into one FINDINGS.md master table: 18 awaiting owner decisions, 4 confirmed designed, 4 flaky E2E failures that didn't reproduce, 1 in progress. Consolidation exposed shared root causes and enabled risk-based ordering where impact, likelihood, and cheap fix beat ticket age. Work isn't closing tickets, it's ensuring every finding has an owner, a decision, and an order.
Four QA session reports lay open on screen, each carrying different finding codes: R4-1, R4-5, R5-2, R3-1. All were born from sessions that worked the same way: a browser driven through MCP with accessibility snapshots [3], so every finding arrived with its dowry of screenshot, SQL row, and network log. All looked urgent, all came from different modules, and I had to answer one question: which gets fixed first?
The key lived in a single commit containing just 64 lines of markdown, FINDINGS.md. It turned four scattered reports into one master table, and the numbers spoke: 18 findings waiting on an owner decision, 4 resolved or confirmed as designed, 4 baseline E2E failures that didn't reproduce on re-run, and 1 in progress. The list healed nothing, but it changed the question from "what is broken?" to "who decides, and when?".
A Finding Without a Decision Is Just a Note
Piled-up tickets create an illusion of productivity: the count drops every day, the direction never gets clearer. The four baseline failures marked "not reproduced" are the most honest example. Six E2E failures from the early phase were re-run one by one: services passed 6/6, maps 5/5, portal 8/8, admin CRUD 5/5 solo. Nothing reproduced. But nothing got fixed either, because nothing known-to-be-broken was found. "Not reproduced" is a status, not a fix; conflating it with "done" is the fastest way to close a problem nobody understands. Flakes are a strange disease anyway: half of it is about interaction contracts nobody reads, like dialogs being auto-dismissed when no handler is registered [1], and half is environments drifting without anyone writing it down.
Consolidation also forces root causes to meet. Phone format validation is absent in two modules at once: R4-1 in informasi and R4-7 in wisata, different layers (announcementadmin vs entityadmin), the same question. Because they were merged, one decision closes two modules, with a shared validator in a common directory as the candidate solution. Media orphans work the same way: R4-5 started in announcement, but permanently deleting a wisata entity leaves the exact same trace, a database row and a physical file that didn't get deleted. Two modules, one decision. Page-detail double-encoding (R3-1 and R3-2) practically found itself during the root-cause investigation: slugs with commas and apostrophes, Next.js params getting encoded twice, so treating it as one fix slice instead of two is the logical call.
Order by Risk, Not by Ticket Age
The priority order suggested by the list itself reads like an impact chain: R4-2 with high severity comes first, the "ALL" scope treated literally so bidang-tagged content disappears from both admin lists and public pages; then R4-4, an API error silently logging admins out because the token gets purged without telling a 401 from a 5xx; then the two old security findings, dev credentials in a repo and a db-push that overwrites users. Everything else is display and polish: the wrong-language toast, the silent upload, the healthcheck calling a tool that doesn't exist in the image.
One category never entered the count at all: findings that turned out not to be findings. Review moderation was suspected as a bug because reviews published instantly without a queue, until the code was traced and the design confirmed: instant publish, moderation moving after the fact, not before. The master list keeps it as a separate clarification so the next session doesn't repeat the same investigation. A wrong conclusion costs more than a correct finding, because it eats someone's time twice.
It's also worth noticing that some findings on the list are old. Dev credentials in the open repo date back to the handover; the healthcheck using a tool missing from the dev image is equally old. A finding's age doesn't raise its priority by itself; what raises it is the combination of impact likelihood and cheapness of the fix. An old finding with a one-line fix is actually a good candidate to finish early, before the owner burns out staring at the list.
The filtering question I used afterwards is simple: if this fix waits another week, does data get lost or do users get misled? If not, it can wait its turn. If yes, it cuts the line. Severity helps, but that question is what turns numbers in a table into decisions.
My whole view of the work shifted once FINDINGS.md existed: my job isn't closing as many tickets as possible, it's making sure every finding has an owner, a decision, and an order. Execution speed without direction just accelerates being wrong.