A Census of 348 SIPP Endpoints Before the Scraper
Building scrapers from endpoint guesses fails fast: a census of 348 SIPP endpoints and 12 official sources sets the foundation before the first crawler runs.
TL;DR
A census of 350 Indonesian district courts found 348 active SIPP endpoints, confirming no central API exists. Twelve other official sources were tiered by verification, with the LKPP blacklist pilot first, then a phased per-court SIPP crawl. The real value comes from reading all sources as one curve covering pre-distress, distress, and post-collapse signals.
The first fetch script was pointed at the case portal of an Indonesian district court, and the screen showed no case list. What appeared instead was a single reCAPTCHA page. Before that moment, planning for the signal-intelligence product still rested on a spreadsheet of endpoint guesses plus one big assumption: surely there is one official central API for searching cases across all Indonesian district courts.
That assumption was wrong. Official Supreme Court documentation describes SIPP, the Case Tracking Information System, as a web application that runs the case-management business process digitally at each court office [1]. There is no single search gateway for district-court cases. Every court operates its own SIPP installation, with versions and configurations that can differ from one another. The signal data exists, but it is scattered across hundreds of separate doors.
The reCAPTCHA page itself was a reminder that court servers carry anti-automation layers. Respectful data collection is therefore not a style choice but an absolute constraint: randomized delays between requests, an honest User-Agent, aggressive caching, and public data only.
Census result: 348 of 350
A census of 350 district courts, completed by DemandScope in early October 2026, found 348 active SIPP endpoints. Of those, 322 follow the canonical sipp.pn-*.go.id pattern, 24 use naming variants, and 2 show no detectable public surface. All of them group under 30 High Courts. Two courts without endpoints are not a bookkeeping failure; the gap is itself data, because endpoint existence cannot be assumed, only verified one by one.
Beyond the courts, 12 other official sources entered the registry under three verification tiers. Tier A means the page was actually opened during the inventory and its numbers confirmed. Tier B means existence is confirmed through search evidence, but the data structure is untested. Tier C is a doctrine lead only. An example of tier C: the tax-collection chain, where no public surface has been found so far, so its existence must not be claimed.
ROI scores and execution order
Every source gets an effort score and a value score on a 1-to-5 scale. The LKPP National Blacklist sits at effort 1, value 4. The LKPP PPID page describes it as the system holding the identities of suppliers sanctioned with a blacklist by procurement officials, under Perpres 16/2018 as its legal basis [3]. Its structure is the cleanest, its cost the lowest, and its output is immediately usable for due diligence.
The Supreme Court's Direktori Putusan sits at effort 2, value 5 because it is centralized. When checked in early October 2026, its public counters showed 1,282,151 civil rulings, 56,134 special-civil rulings including bankruptcy and PKPU cases, and 472,534 uploads during 2026 [2]. Those counters move every day, so they must be recorded as a dated snapshot, never as a permanent total. Per-court SIPP sits at effort 4, value 5: bankruptcy petitions appear there earliest, and maintaining a crawl of 348 endpoints is an aggregation moat that is hard to copy.
Execution follows those scores. Run the pilot against the LKPP blacklist first to prove the data pipeline end to end with the cheapest source. Continue with probes of tier-B sources to map their structures. The SIPP crawl starts after that, phased court by court, beginning with the commercial courts in Jakarta.
The signal curve: before, during, and after
The product value only appears when these sources are read as one curve, not as separate tables. The pre-distress phase emits procurement-sanction and tax-collection signals. The distress phase surfaces bankruptcy or PKPU petitions in SIPP along with court rulings. The post-collapse phase fills with mandatory curator announcements: Article 15(4) of Law 37/2004 obliges the curator to announce a summary of the bankruptcy ruling in the State Gazette and at least two daily newspapers, no later than five days after the ruling is received [4]. A company crossing this curve leaves traces in three different source clusters, and only a system that maintains all three can see the full line.
One status note: Law 37/2004 remains in force, but several of its articles carry conditionality from Constitutional Court rulings, so any citation of it should follow its official status page [5].
The decision that came out of this census is concrete. The LKPP blacklist pilot enters the next sprint, tier-B sources enter the probe queue, and the SIPP crawl is phased by High Court jurisdiction. One assumption, meanwhile, is off the table for good: the existence of a central API.