Through One Door: A Retry-Safe Harvest Ingest Contract
A scraping workspace kept apart from the app, and a replayable single-door API delivery: resumable queue, stable dedupe key, ship log.
TL;DR
A mid-run interruption exposed the risk of writing scrape results straight into the app database, so this commit isolates all scraping tools and stores per-page, per-subtab queue rows with status and attempt logs on disk. Resuming is then deterministic. Shipping happens separately through one ingest endpoint, using stable deduplication keys and exponential backoff so idempotent retries make replays safe.
The harvest stopped at 18 of 70 cases. Another 52 AJAX sub-tabs were still waiting when the session had to be cut short. Under the old layout, scraping scripts, raw captures, and test fixtures lived side by side with the main application code; a single bad import was enough to touch application storage by accident.
The tempting assumption: write harvest results straight into the application database as collection proceeds. One pass, data immediately available, job done. The weakness only shows when the session dies mid-run. The collection lifecycle is chained to the target database's availability, and an interrupted run leaves a question nobody can answer: which parts made it in, and which did not.
This commit takes the opposite road. Every scraping tool moved into a separate workspace directory under one strict doctrine: nothing inside it may write application storage directly. Harvested data enters only through one door, an API ingest endpoint, via a separate shipping step that can be replayed.
A Resumable Work Queue
The sipp_queue.py module breaks the work into small rows. One detail page is a row; each AJAX sub-tab inside it, from rulings to hearing history, is another row. Every row stores a pending -> ok | error status plus attempt counts and failure notes.
A status report answers three questions exactly after any interruption: what is done, what is pending, and which URL misbehaved. The URL log records every page actually opened with its outcome, separate from the work queue and from the event log that serves as the audit trail. Because all state lives on disk, a broken run resumes without guesswork.
The One-Way Shipping Bridge
The ship.py script closes the chain. It wraps harvested data into items and posts them to the endpoint one by one. Its contract is strict: every item carries a stable deduplication key built from the f"{source_type}:{source_id}" pattern, pairing the source-type identifier with the case number within that source. The same key always names the same logical datum, on the first shipment and the fifth.
On the client side, ship-log.jsonl records every attempt with its outcome. On a re-run, items without a recorded success ship again, while items already logged as ok or skipped-duplicate are skipped. On the server side, the same key drives an upsert instead of a duplicate. The contract is idempotent at both ends.
The full script carries extra details such as a dry-run mode and a path fallback for relocated endpoints. The core shipping pattern condenses to this.
import time
def ship_with_retry(dedupe_key, payload, max_attempts=3):
for attempt in range(1, max_attempts + 1):
status = post_item(dedupe_key, payload)
if status in ("accepted", "duplicate"):
record(dedupe_key, status)
return True
time.sleep(2 ** attempt)
record(dedupe_key, "error")
return False
Transient failures do not trigger request storms. The delay between attempts grows exponentially, keeping the load toward the server under control and giving the target system time to recover.
The Contract That Makes Retries Safe
Retry safety here is not a bolted-on trick; it is a definition HTTP already carries. A method is idempotent when several identical requests have the same effect as one [1]. Clients may automatically repeat a request after a communication failure because the outcome is known to be unchanged, and even POST may be retried when the endpoint is designed to be safe for it.
Data shipping between systems almost always runs with at-least-once guarantees: a message can arrive more than once. End-to-end exactly-once is impractical to guarantee, and the durable answer is a receiver that lets duplicates pass without effect [2]. The requirement is one line long: the deduplication key must identify the same logical message uniquely and consistently across every redelivery.
In practice, the server still has to enforce idempotent semantics, since the specification does not force endpoints to behave according to their method's definition [3]. The operations side states the same principle: APIs with side effects should be designed idempotent so retries are safe, while growing delays keep load even while the system recovers [4].
The easiest part to miss in this commit is not the shipping code; it is the boundary. Once harvested data must pass through one door under a replayable contract, where the API runs becomes a configuration matter instead of a scraper matter: local today, another server six months from now, without touching the harvest code again.