The Raw-First Pattern for Gated Data Harvests
Why gated data harvests save verbatim raw text before parsing: an idempotent ledger, a resumable cursor, and retroactive rebuilds.
TL;DR
When scraping captcha-gated, session-locked court portals, save the verbatim rendered text first and parse later from local files. Resumable cursors and idempotent upserts keep crashes from corrupting data. This bronze-layer approach means parser fixes apply retroactively without ever hitting the fragile site again.
The harvest session stops abruptly on page eight of a twenty-page case list. Fourteen case detail pages have been captured one by one through an interactive browser preview pane. In this restricted environment, standard click and open actions are disabled. Navigation relies entirely on pressing the Enter key to advance through the interface. The source site is a regional court portal bound by strict session cookies and protected by captcha interstitials. Re-accessing these specific pages later is never guaranteed.
When building an ingestion pipeline for such restrictive environments, the immediate instinct is to parse the data during the capture phase. Extracting structured fields on the fly feels computationally cheaper. The logic suggests a single pass where the document is downloaded, parsed, and inserted into a database simultaneously. This approach assumes the parsing logic will be perfect on the first attempt and that the source will remain perfectly stable.
In practice, parsing errors are a certainty, while source re-access remains a risk. The reliable approach saves the verbatim rendered text first, deferring structured extraction to a secondary step. The tool writes raw text files to establish verbatim provenance, capturing information about the entities and activities involved in producing the data to assess its reliability [1]. Structured parsing runs afterward against these local files. A dedicated rebuild subcommand can regenerate the ledger and flat tables directly from the raw text. This means parser fixes apply retroactively to already-harvested data without needing to hit the gated site again. This raw-first methodology converts a network risk into a local compute task. The initial capture focuses solely on survival, ensuring the text is secured behind the captcha gate before any fragile regular expressions attempt to dissect it.
Idempotent Bookkeeping and Resumable State
To support deferred parsing, the bookkeeping script relies on a resumable cursor stored in state.json. This file tracks the run identifier, the current list page, pending references, and the set of completed items. Because the interactive session can crash or pause at any moment, the next run simply resumes from this exact cursor. Actions are logged in run-log.jsonl, where each structured record is processed one line at a time, making it highly effective for continuous log files that work seamlessly with standard Unix text tools [4]. The final structured output lands in ledger.jsonl, which remains idempotent by case number. Using an upsert mechanism, the database insert behaves as an update or a no-op if a uniqueness constraint is violated, automatically skipping duplicates [2]. This guarantees that interrupting the harvest midway never corrupts the existing dataset.
The Bronze Layer at a Single-File Scale
This workflow mirrors the medallion architecture used in large-scale data engineering, specifically the bronze layer. The bronze layer acts as an as-is landing zone designed for historical archiving, data lineage, and auditability [3]. It allows for reprocessing without rereading the source system. By treating the local text files as the bronze layer, the harvest isolates the fragile network interaction from the fragile parsing logic. The raw files become the absolute source of truth.
The same trade-off has a boundary. An open API with stable endpoints can be re-fetched any time, so keeping verbatim captures is a bonus rather than a requirement. The pattern pays for itself exactly when re-access is expensive: captcha-gated sources, expiring sessions, or pages that change between two visits.
If a date format changes on the source site, or a new HTML wrapper breaks the extraction regex, the pipeline does not need to request the pages again. The parser is updated locally, the rebuild command is executed, and the structured output is corrected using the preserved raw text. The network boundary is crossed exactly once per document. Six months later, when the parser changes again, the steps stay the same while the raw files remain untouched.