Skip to content

Tracing Harvested Data with Agent Guardrails and SQLite Logs

Adityo Guni Waluyo

Two local guards for scraping pipelines: an AGENTS.md rulebook for coding agents and a per-target SQLite audit log made idempotent with UPSERT.

TL;DR

When a harvest run died, git couldn't trace which commit produced the data—it records file changes, not execution context. The fix was two guards: an AGENTS.md rulebook forbidding writes to harvested data and secrets in commits, plus a per-target SQLite audit log. UPSERT on a unique commit-case-tab key keeps re-runs idempotent, turning provenance recovery into a single query.

Mid-run, the harvest session on the DemandScope project died. The terminal showed a process that had stopped half-done, and one question sat there without an answer: which commit carries this case's data? My first guess was that git history would be enough to trace it.

That assumption was wrong. Git only records that files changed, not which execution produced which rows. A single commit can bundle output from several interrupted runs, and the link between execution context and the data on disk simply was not written down anywhere.

The fix that lasted was not a fancier tracker. It was two local guards: a rulebook the coding agents must obey, and a per-target audit log in SQLite.

The Rulebook

AGENTS.md is an open format now used by more than 60 thousand open-source projects [3]. An analysis of over 2,500 public agents.md files found a clear pattern in the effective ones: a specific persona, executable commands placed early, code examples instead of long prose, and explicit never-touch boundaries [5].

For DemandScope the two hard data rules are: agents never write into the harvested-data directory, and agents never carry secrets into a commit. Commit discipline lives in the same file. Every harvest commit must be recorded in the audit log.

Codex loads these files before doing any work. Discovery walks from the global home directory to the project root and then into nested directories, files closer to the working directory override earlier ones, empty files are skipped, and the combined prompt is capped at 32 KiB by default [6]. A rulebook has to stay short; whatever lands past the cap silently never loads.

The Audit Log

The second guard is one SQLite database file per scraping target, next to that target's raw data. Every harvested unit, a commit-case-tab combination, gets one row. UPSERT is what makes re-runs idempotent: an INSERT that becomes an UPDATE or a no-op when it would violate a uniqueness constraint, and it only fires on uniqueness constraints [2]. The guarantee is a schema decision, so the UNIQUE key has to be declared up front:

CREATE TABLE harvest_log (
    commit_hash TEXT NOT NULL,
    case_no     TEXT NOT NULL,
    tab         TEXT NOT NULL,
    status      TEXT NOT NULL,
    ts          TEXT DEFAULT (datetime('now')),
    UNIQUE(commit_hash, case_no, tab)
);

INSERT INTO harvest_log (commit_hash, case_no, tab, status)
VALUES ('a1b2c3d', '2026/001', 'putusan', 'ok')
ON CONFLICT(commit_hash, case_no, tab)
DO UPDATE SET status = excluded.status, ts = excluded.ts;

SQLite fits this job by design. It competes with fopen(), not with client/server databases, and the official docs treat sites under 100K hits per day as comfortably within range [1]. An audit table is nowhere near those limits.

Each row is also miniature provenance: information about the entities and activities that produced a piece of data, which is what you use to judge its quality and trustworthiness [4]. Formal run trackers ask for more. OpenLineage requires a unique run id plus start and completion-or-failure events for every run [7]; this small log mirrors that contract inside a single file.

SELECT commit_hash, status, ts
FROM harvest_log
WHERE case_no = '2026/001'
ORDER BY ts DESC
LIMIT 1;

When the next interruption happens, that query is the whole investigation. Every row now has a verifiable origin, and the answer no longer depends on reconstructing a session from memory.

## Sources [1] Appropriate Uses For SQLite — sqlite.org [2] UPSERT — sqlite.org [3] AGENTS.md [4] PROV-Overview — W3C [5] How to write a great agents.md — GitHub Blog [6] Custom instructions with AGENTS.md — OpenAI Developers [7] OpenLineage Spec

Related articles