Skip to content

The Morning My Research Tool Swallowed Its Own Medicine

Adityo Guni Waluyo

One pruned line in the morning build settled it: an unpruned vector index answers with documents that no longer exist, and free-form data is a contract.

TL;DR

Renaming a research file left a stale embedding in a JSON vector index that kept returning a ghost document. The fix added tolerant reading for summary field aliases, automatic pruning of missing files during builds, and a self-testing suite. It is a sixty-line reminder that small local tools need honest indexes and defensive boundaries.

One line from the morning build stopped me mid-sip: pruned (file missing/renamed). An entry got dropped from my local vector index because its file was gone. I had renamed a research dossier file the night before and forgot one thing: the JSON file holding its embedding had no way of knowing. A tool I wrote for my own use had just caught its own boomerang from the day before.

My first guess was dismissive. A frontmatter field alias? Tiny fix. Pruning stale index entries? Housekeeping to keep the JSON file from getting fat. That guess lasted until I pictured the alternative: the research search query keeps returning the renamed dossier, its cosine score still high, and the article cron's knowledge check trusts a document that no longer exists. An unpruned index isn't just wasteful. It answers with files that physically aren't there anymore. This was about the truthfulness of query results, not tidiness.

The fix landed as +64/-4 lines in one Python file, but it carries three defense layers I now consider mandatory for any small tool: tolerant reading at the boundary, index hygiene, and tests that can prove themselves wrong.

Tolerant Reading at One Boundary

The root cause was mundane. My older research dossiers were written through an agent using an Indonesian convention, key ringkasan:. The new ones are in English, key summary:. Two writers, two conventions, same human. This is unavoidable, and [2] Hyrum's Law sums up why: with enough users, every observable behavior of a system gets depended on by somebody, no matter what the contract promises. At my scale, that "somebody" is me three months from now, wondering why the summaries came back empty.

The fix is one function, summary_field(), as the single door for reading the summary field, with layered fallback:

# one read boundary: every caller goes through this, nobody reads raw frontmatter
FIELD_ALIASES = {"ringkasan": ("ringkasan", "summary")}

def summary_field(path):
    for name in FIELD_ALIASES["ringkasan"]:
        value = frontmatter_field(path, name)
        if value:
            return value
    return ""  # even empty is non-fatal: the L0 fallback grabs the first core sentence

The opposite approach is a strict parser that trips over its own schema. Fowler documented the failure mode in Tolerant Reader: schema-driven binding breaks exactly when the data provider adds a field that shouldn't be breaking, and [1] his recommendation is to be as tolerant as possible when reading data, take what you need, ignore the rest. Postel's law for local data. A bonus of consolidating aliases in one function: the next convention change touches one place instead of scattering through the whole codebase.

Ghost Documents and a Burden with No Engine to Blame

The second layer is dead entries. With a real vector database, the engine handles this. [3] Qdrant implements deletions as soft deletes using a bitmask, the index doesn't rebuild on every delete, and a deleted point becomes instantly inaccessible through the API. [4] Pinecone takes a different route: an upsert with an existing ID overwrites the entire record. My setup is far more plebeian: one JSON file of vectors, cosine similarity computed at query time, no database. The consequence is clear. There's no engine left to blame, so the build layer carries the load.

So cmd_build now checks every entry before using it: does the file still exist? If not, drop it and print the line. That was the line from this morning. Six lines of core logic, and those six lines made my index honest again.

A Selftest Deliberately Built to Be Wrong

The third layer is the one I think most people skip for personal tools: a test suite. The commit added a selftest subcommand with small fixtures. Aliases must resolve. An empty field must fall through to the fallback instead of crashing. The Findings section must parse for the L0 fallback. My favorite part is the deliberately wrong assertion: a comparison against a clearly different string, which must fail if the implementation breaks. It's mutation testing in miniature. A gauge that can't possibly be wrong is a broken gauge, and a suite that stays green proves nothing except that it isn't checking anything.

My position now is firm: free-form local data is a false promise. A format that's "clear enough for me" turns into an implicit contract the moment there are two writers, two months, two conventions. I'd rather pay 64 lines now than debug a knowledge check that trusts ghost documents six months from now. Tomorrow morning, if another rename happens, that pruned line will show up again, and it won't be an error. It's the tool admitting it isn't perfect, which is exactly why I can trust it.

Sources

Related articles