Skip to content

Recovery-Only Rebuild: Lessons from a Corrupted Ledger

Adityo Guni Waluyo

A rebuild turned 20 healthy ledger rows into 77 UNKNOWN. The fix: a recovery-only contract that fills gaps and never rewrites healthy rows.

TL;DR

A rebuild tool corrupted a court-rulings ledger, inflating twenty healthy rows to seventy-seven, mostly tagged UNKNOWN. The cause was structural: per-tab raw files re-parsed as new records, proving raw-is-truth doesn't make re-parses safe. The redesign makes rebuild recovery-only: healthy rows stay verbatim, only missing cases are appended, sub-tab files are skipped, and any UNKNOWN aborts all writes.

I ran rebuild_ledger against the ledger of a court-rulings portal on the DemandScope project and watched twenty healthy rows turn into seventy-seven, fifty-seven of them tagged UNKNOWN. Recovery was a git checkout of the ledger file. The tool meant to maintain the data had just rewritten it into garbage.

The cause was structural. Sub-tab raw files, fragments captured per tab rather than per case, fed the detail parser and came out as new records. The old docstring promised raw is truth, and I read that as a license: if the source is unchanged, a full re-parse must be safe.

Why the Assumption Fails

Raw-is-truth gets the hierarchy right and the safety model wrong. Knowing the raw files are intact says nothing about what a re-parse is allowed to replace. Storage engines make the same distinction. SQLite database files resist corruption, partially written transactions roll back automatically after a crash, yet they remain plain files that any process can overwrite, and the library cannot defend them [1]. My rebuild script was exactly such a process.

Vendors treat recovery paths the same way. PostgreSQL documents pg_resetwal as a tool to use only as a last resort when the server will not start due to corruption; it demands the explicit force flag, and after running it you should immediately dump and restore, with no data-modifying operations before the dump [2]. The SQLite CLI ships .recover to recover as much data as possible from a corrupt database, wording that promises best-effort, not fidelity [3]. A homegrown rebuild deserved the same guardrails, and it had none.

The Recovery-Only Contract

The redesign narrowed what rebuild is allowed to do. Four guards now define it:

  • Healthy rows stay verbatim, matched by case number.
  • Only rows genuinely missing from the ledger get appended.
  • Sub-tab files are ignored entirely.
  • If any new row parses as UNKNOWN, nothing is written at all.
def rebuild_ledger(raw_dir, current_ledger):
    new_rows = []
    for f in raw_files(raw_dir):
        if f.name.endswith("-tab.txt"):
            continue  # sub-tab files are not detail rows
        row = parse_detail(f.read_text())
        if row.status == "UNKNOWN":
            abort("nothing written")  # all-or-nothing
        if not ledger_has(current_ledger, row.case_no):
            new_rows.append(row)
    current_ledger.extend(new_rows)  # healthy rows stay verbatim

The bigger change was not in the code. It was in when the tool may run at all. Rebuild stopped being a refresh button and became a recovery tool, called only when a verified gap exists.

The Append-Only Foundation

Recovery-only works because the raw layer never changes. Harvest sessions append; they never edit history. Kafka applies the same philosophy at scale: events are stored durably and are not deleted after consumption, with retention configured per topic [4]. Immutable raw files are what make a conservative rebuild possible in the first place.

Two details carry the rest. For the SQLite-shaped parts of the pipeline, idempotent writes only exist when UPSERT fires on a declared UNIQUE key [5]; uniqueness is a schema decision, not a query trick. And the journal mode that governs transaction handling stays at its default of DELETE [6], one more moving part I deliberately did not touch mid-repair.

What is healthy does not need rebuilding. That sentence is now the tool's contract, and the incident that taught it is on record in four guards that never write a byte they cannot justify.

## Sources [1] How To Corrupt An SQLite Database File — sqlite.org [2] pg_resetwal — PostgreSQL Documentation [3] Command Line Shell For SQLite (.recover) — sqlite.org [4] Apache Kafka Documentation [5] UPSERT — sqlite.org [6] PRAGMA Statements — sqlite.org

Related articles