Skip to content

Review Status Tokens: Making Review Debt Grepable

Adityo Guni Waluyo

Four review:* tokens plus a board validator make agent review debt visible, countable, and batchable, while security surfaces never wait in a queue.

TL;DR

Reviewing every tiny slice cost 50-60k subagent tokens per round regardless of diff size, making review the largest session expense. The fix: every done row must carry a review token, enforced by a board validator, with non-security work batched until the queue hits three or the session closes. Security surfaces skip the queue, and tokens make review debt greppable.

I closed the last done-row of a work session on KotaPortal and the review queue had quietly swollen. My first assumption: reviewing every tiny slice individually is the safest path. More review rounds, fewer bugs slipping through.

It turns out that instinct is expensive. One review round per slice costs roughly 50-60k subagent tokens (project worklog), no matter the diff size. A two-line change pays the same as a two-hundred-line one. When a single session births a dozen small slices, review becomes the largest line item in the session's budget, and nobody records it piling up.

Four Tokens and One Validator

The fix is small in shape. Every done row on the board must now end with one of four tokens: review:pending, review:code, review:code+sec, or review:none. A board validator rejects any done row without a token, the DONE-REVIEW invariant (project worklog). The principle: done without review is allowed, as long as it is recorded.

The policy is adaptive, not binary. Work on the trigger list (auth, sessions, RBAC, upload surfaces, injection, API contract files, destructive migrations, security fixes) is still reviewed immediately, never batched. Everything else closes as review:pending and waits for one combined code-only review when the queue reaches three or more, or when the session closes.

A Time Fence, Not a Magic Number

The threshold of three is a fence, not a prayer. Google's engineering practices name one business day as the maximum response time for a review request, and most complaints about review trace back to a slow process [2]. Batching without a time bound just moves the problem: feedback arrives after the context has gone cold. A measurable trigger cuts cost without turning debt into stale work.

OWASP places the review process for code and configuration changes among its integrity controls, not ceremony [4]. That is exactly why security surfaces leave the batch from second one: A08 is about assumptions taken without verifying integrity.

Debt You Can Grep

The most useful side effect: the board became machine-auditable. grep -c "review:pending" returns the queue length in one command. The philosophy mirrors Conventional Commits, which made commit history machine-readable for tooling [3]: review status now has the same shape, visible and searchable instead of hidden in someone's head.

Backfilling 16 legacy rows proved the pattern: 2 code+sec, 5 code, 9 none, 0 pending (project worklog). But the biggest lesson was not the numbers. Expectations written at card level did not map one-to-one onto board rows, because several cards lived only on the board and had no row to carry a token. Rollups must live at one level. Mix card and row accounting and expectations break silently, with the deviation surfacing only at execution time.

Reviewer speed is part of the design too. Industry practice treats small, independent changes as the unit of review, with fast responses as the backbone of team health [1]. Status tokens guarantee that "not reviewed yet" is a visible state, not an ambiguity left to fester until someone asks why this bug shipped.

What surprised me most was the mental load. Before tokens, review status lived in session memory: which slices were checked, which only pretended to be. Every time I resumed work I opened the board and hesitated. With tokens the hesitation is gone, because the answer sits on the row itself, not in my recollection.

If you want to adopt this, start with the validator, not the policy. Without an invariant rejecting tokenless rows, batching is a promise in a document. With one, the policy becomes a mechanical consequence: an invalid row can never count as done.

One note on the number: the threshold of three was chosen from the context size of a working session, not from a benchmark. The knob belongs to the project owner and may move anytime. The principle may not: a review queue needs a measurable trigger, and security surfaces stay out of the queue.

Sources