Skip to content

A Daily Signal Ingest Pilot from Public Registries

Adityo Guni Waluyo

Two registries passed probing: scrapers behind one seam, idempotency forced via content hash, and a daily 05:00 ingest with jitter.

TL;DR

A pilot ingests daily signals from the LKPP procurement blacklist and three courts' SIPP bankruptcy/PKPU case lists through a single SignalSource interface. Idempotency comes from content hashes and unique constraints, so re-runs never duplicate data, and one court failing doesn't kill the run. The job runs once daily via APScheduler under strict politeness rules, with privacy safeguards built in.

The pilot plan called for daily ingest from two public registries: the government procurement blacklist and the bankruptcy and PKPU case lists of district courts. One decision shapes the entire system: the same ingestion will run every morning, over and over, without ever duplicating a signal that is already in.

The quickest idea is a plain cron that inserts every fetched row into one table. That plan survives exactly one count: the blacklist page carries dozens of sanction rows that shift every day, and the same rows come back tomorrow morning. Without a mechanism, the table fills with duplicates within a week.

So the principle is one: every ingest run must be safe to re-run at any time. The word is idempotent, and the whole design below bows to it.

Pick the two sources that passed probing

The first source is the National Blacklist maintained by LKPP, the information system holding the identities of suppliers sanctioned by procurement officials, with Perpres 16/2018 as its legal basis [1]. The second is the SIPP case-list pages of three courts still open to measured access. SIPP is officially a per-court web application rather than one central portal [5], so each court is configured on its own.

The two sources behave differently. The blacklist's first page renders server-side with the newest sanctions on top, while deeper pages are JavaScript-driven; a daily snapshot of page one is enough for the pilot. A court's case list is a table with case number, classification, parties, and status columns. The per-case-type menu turns out not to filter anything, so classification happens locally: the Pdt.Sus-Pailit and Pdt.Sus-PKPU case-number patterns mark bankruptcy and PKPU signals, while company names come from the parties column.

Put every scraper behind one seam

Both sources are implemented as scrapers that satisfy the same interface. That interface is SignalSource: other components never need to know whether a signal came from the blacklist or from a court. Normalized results land in two tables:

class SignalSource:
    source_id: str

    def fetch(self) -> list[dict]: ...

# signals: id, source_id, source_ref, company_name_raw,
#   company_name_normalized, signal_type, event_date,
#   details (JSONB), content_hash
# ingest_runs: id, source_id, started_at, finished_at,
#   status, pages, items_seen, items_new, error

Idempotency is enforced at the database level, not only in the design doc. A content_hash column is computed from the signal's content, and a unique constraint on the source_id and content_hash pair rejects rows that were already ingested. The second day's run reads the same page, computes the same hash, and the database refuses it. Re-run three times a day and the result stays one row per signal.

Partial failure is part of the design too. One court being down must not fail the whole run: the process moves on to the next source, and the failure status lands in ingest_runs. For reading results, a tenant-scoped API serves signals plus a daily digest grouped by company.

Schedule once a day, pace every request

The ingest schedule is once a day around 05:00 Western Indonesia Time, run from inside the application container through the Python scheduler rather than system cron. APScheduler ships a CronTrigger with timezone and jitter parameters [2]:

from apscheduler.triggers.cron import CronTrigger

trigger = CronTrigger(
    hour=5, minute=0,
    timezone="Asia/Jakarta",
    jitter=60,
)
# setara crontab: 0 5 * * * (Asia/Jakarta)

Inside a run, politeness holds as hard rules: the User-Agent names the service, 2 to 5 seconds of jitter between requests, a hard cap of 12 requests per run, raw HTML cached before parsing, and one pass per day. RFC 9309 states that robots rules are not a form of access authorization [3]; politeness is part of the design from the start, not a reaction to a warning.

Gated sources are recorded, not fought. No captcha solving, no WAF evasion; a court that refuses is dropped from the configuration until it becomes reachable again. One privacy note travels with the data: case rows may carry director names, and Law 27/2022 on Personal Data Protection governs personal-data processing [4], so anonymization sits on the design-consideration list from the pilot phase.

The pilot ends when one sentence can be demonstrated on real data: show last week's signals for one company. Two scrapers, one seam, one shared schedule. The rest is execution.

Sources

  1. PPID LKPP: National Blacklist Application
  2. APScheduler Documentation: apscheduler.triggers.cron
  3. RFC 9309: Robots Exclusion Protocol
  4. JDIH BPK: Law No. 27 of 2022 on Personal Data Protection
  5. Supreme Court Documentation: Introduction to SIPP Backup

Related articles