GraphQL from the Site's Own Bundle: 7 Requests over 533 Pages
Finding a site's own GraphQL endpoint in its public JS bundle: 5330 rows in 7 requests, not 533 HTML pages.
TL;DR
Expecting a tedious HTML scraping job for Indonesia's procurement blacklist, the author found an unauthenticated GraphQL endpoint exposed by the site's XLSX export. A perPage of 800 meant the full 5,330-row dataset downloaded in just seven requests and about ten seconds. The daily snapshot plan was retired in favor of a simple full-dataset sync.
GetBlacklists, fired straight at /graphql [10]. And the best hint came from the site's own XLSX export button, which happily revealed the sanctioned request shape including perPage: 800.
That fits how the framework works. Next.js renders layouts and pages as Server Components by default, but client-side apps lean on Client Components for state and browser-side data fetching [8]. With no auth, captcha, or cookies in the way, one request with 800 per page returned page 1 of 7: 5330 rows in total.
The full sweep took 7 requests, roughly 10 seconds, with a bot User-Agent paced at one request per second. The whole dataset in one command instead of 533 separate HTML pages. And the dataset is far richer than the visible table: masked NPWP taxpayer numbers, provider addresses, pagu/hps contract values, violation regulation text, LPSE/KLDI/satker origin, SK numbers, publish and expiry timestamps.
From my recount of the committed 7.7 MB JSON: 5330 rows, with 348 PUBLISHED, 4916 EXPIRED, 63 CANCELED, and 3 CANCELED_TEMPORARY. Masked NPWP numbers appear on 4169 rows. Tender pagu or hps values exist on 3006 rows, about 56 percent. Start dates span 1905, one bogus row, to 2026. 27 rows start in September-October 2026, exactly the fresh-signal window a daily digest needs. 4260 unique provider names in total.
Offset pagination has known performance and security downsides on large datasets [9]. At 5330 rows this one is small, so seven honest requests beat any cleverness. GraphQL introspection would even map the schema's types and fields without guessing [11].
This discovery retired the daily page-1 snapshot plan entirely. Ingest is now a plain HTTPS GraphQL client, bot User-Agent, one request per second, syncing the full dataset daily in 7 requests and deduplicating by row id. The 5330-row dump is the day-one historical baseline; older rows never get re-fetched. The HTML parsing lane is retired, and the page-1 JSON survives only as lane provenance.
The request pattern is short:
import requests
url = "https://daftar-hitam.example.id/graphql"
headers = {
"User-Agent": "DataBot/1.0 (Research Purpose)",
"Content-Type": "application/json"
}
payload = {
"query": """
query GetBlacklists($page: Int, $perPage: Int) {
blacklists(page: $page, perPage: $perPage) {
data {
id
providerName
status
npwp
paguHps
startDate
endDate
}
pageInfo {
currentPage
totalPages
totalData
}
}
}
""",
"variables": {"page": 1, "perPage": 800}
}
response = requests.post(url, json=payload, headers=headers)
print(response.json())