Not a Race Condition, but a Dual-Write
One ticket skipped two statuses, a citizen got a duplicate email: the root was a dual-write. The fix is a transition map, 409, history, and an outbox.
TL;DR
Duplicate emails weren't a frontend race condition but a dual-write bug where database updates and SMTP sends weren't atomic. The fix uses a transactional outbox, saving the email to an outbox table in the same Postgres transaction for a background worker to deliver. Illegal status jumps are now blocked early with 409 Conflict using a central transition map.
Not a Race Condition, But a Dual-Write: Learning from Duplicate Emails in KotaPortal
Meta (untuk sistem): Experience fixing duplicate email bugs and status jumps in a complaint system using the Transactional Outbox pattern and HTTP 409 Conflict. Slug (untuk sistem): /en/blog/status-transition-409-outbox
I was checking the complaint table in the KotaPortal system, and my eyes caught a weird row. The ticket status changed from "masuk" directly to "selesai", even though the normal flow had to go through "diproses" first. What was worse, the reporting citizen received two notification emails within a one-minute interval for the exact same button click.
An Initial Guess That Missed the Mark
At first, I thought this was a race condition on the frontend. Maybe two admins accidentally clicked the "complete" button at the same time, or a slow network caused the user to spam the click. I even added a loading spinner and disabled the submit button in the UI for a few seconds after being clicked. I also added debouncing to the API call function.
But the problem kept appearing. After I dug into the backend logs and traced the execution path, my guess was completely off. The issue wasn't in the UI, but in how the backend handled status changes and notifications.
The old code updated the status in the database, then immediately called the SMTP email sending function on the next line within the same function. This is a classic **dual-write bug**. We were trying to write to two different systems without a reliable coordination mechanism.
If this email sending process failed or timed out after the database status had already changed, we ended up with a ticket whose status was updated but the citizen got no notice. Conversely, if there was a flawed retry mechanism at the application layer, the email could be sent multiple times while the database status only changed once. We had no guarantee of atomicity between these two operations.
The Dual-Write Trap and the Outbox Solution
The solution wasn't to add complex retry logic or attempt a heavy distributed transaction. The solution was to strictly implement the status transition 409 outbox pattern.
The idea is simple but highly effective. Instead of sending the email directly, we save that email message into a dedicated table in the same database as part of the entity update transaction. A separate process, a background worker, will later pick up that row and actually send it to the SMTP server.
Here is an overview of how that transaction structure should be formed to guarantee consistency:
This way, PostgreSQL guarantees that the three steps above are a single all-or-nothing operation [3]. If any step fails, for instance due to a constraint violation or dropped connection, the entire transaction will be automatically rolled back. No more ghost emails or statuses changing without a clear history trail.
The background worker will run consistently, picking up rows with a pending status, sending them, and then updating their status to sent. If the delivery fails, the worker can mark it as failed and retry later, without ever compromising the integrity of the main data.
BEGIN;
UPDATE tickets SET status = 'diproses' WHERE id = 123;
INSERT INTO ticket_history (ticket_id, old_status, new_status)
VALUES (123, 'masuk', 'diproses');
INSERT INTO email_outbox (ticket_id, event_type, payload, status)
VALUES (123, 'status_changed', '{"to": "[email protected]"}', 'pending');
COMMIT;
Enforcing Transition Boundaries with 409 Conflict
There is one more important detail often overlooked when building a status engine: handling illegal transitions.
In the old code, if someone tried to change the status from "selesai" back to "masuk", the system sometimes returned an HTTP 500 Internal Server Error due to a failing database constraint. Sometimes it even returned an HTTP 400 Bad Request. Both give the wrong signal to the client or monitoring system.
HTTP 400 is for requests where the format or data validation is wrong from the start. HTTP 500 is for truly unexpected errors on the server. However, this request has the correct format, the server is healthy, but the request conflicts with the current state of the resource.
The most accurate semantics for this condition is HTTP 409 Conflict. According to MDN documentation, the 409 code specifically indicates that the request conflicts with the current state of the target resource [1]. This is the most honest way to say "this transition is not legal from the ticket's current status".
The implementation is simply validating an explicit transition map before executing the database transaction block above. The map itself is a small dictionary: status to list of allowed next statuses, with "completed" mapping to an empty list.
Validasinya cuma satu langkah sebelum blok transaksi: baca status tiket sekarang, cek pasangan old-to-new di peta transisi. Pasangan yang nggak ada di peta ditolak dengan 409 sebelum menyentuh database, tanpa constraint yang harus gagal dulu di dalam transaksi.
I personally prefer the explicit transition map approach above over relying on scattered if-else chains across various service functions. if-else chains are fragile and hard to trace. When a new status is added to the system, a developer might forget to update one of the logic branches, and bugs silently slip into production.
With a centralized transition map, we have **a single source of truth** that is easy to audit. If there is a transition request not on the list, the system immediately rejects it with a 409 Conflict before touching the database at all. This saves resources and keeps business logic clean.
Combining strict transition validation with the outbox pattern makes the complaint system much more resilient. We no longer have to worry about lost emails or statuses jumping around without a clear reason. The boundary between valid and invalid states becomes firm, and side effects like email delivery are handled in a reliable way.
Strict transition validation plus the outbox pattern made the complaint system far more resilient. But one honest note about this pattern: the relay can send the same email twice, for example when it crashes right after sending but before marking the row done. The receiver must be idempotent, usually by tracking the IDs of messages it has already processed [2]. The duplicate email I chased at the start of this article has a twin on the worker side too.
Sources