open-data
Technical notes on web development, DevOps, and AI integration.
3 articles
- 03:35backend
Three Gates of Public Registries Before Writing a Scraper
Probing nine public-registry hosts: two open, one captcha, two WAF. Learn the three access-gate patterns before writing a single line of scraper.
TL;DR: Read-only probes of Jakarta court SIPP portals found mixed access: two open, one captcha-gated, two WAF-blocked, others thin or refusing. SIPP is built per court, so variation is expected; each host gets one polite probe saved as a local fixture. No captcha solving or header tricks; three hosts form the pilot, classified locally since menu tokens don't filter.
#scraping#open-data#sipp - 03:26backend
A Daily Signal Ingest Pilot from Public Registries
Two registries passed probing: scrapers behind one seam, idempotency forced via content hash, and a daily 05:00 ingest with jitter.
TL;DR: A pilot ingests daily signals from the LKPP procurement blacklist and three courts' SIPP bankruptcy/PKPU case lists through a single SignalSource interface. Idempotency comes from content hashes and unique constraints, so re-runs never duplicate data, and one court failing doesn't kill the run. The job runs once daily via APScheduler under strict politeness rules, with privacy safeguards built in.
#scraping#open-data#automation - 02:49backend
A Census of 348 SIPP Endpoints Before the Scraper
Building scrapers from endpoint guesses fails fast: a census of 348 SIPP endpoints and 12 official sources sets the foundation before the first crawler runs.
TL;DR: A census of 350 Indonesian district courts found 348 active SIPP endpoints, confirming no central API exists. Twelve other official sources were tiered by verification, with the LKPP blacklist pilot first, then a phased per-court SIPP crawl. The real value comes from reading all sources as one curve covering pre-distress, distress, and post-collapse signals.
#scraping#open-data#sipp