Skip to content

GraphQL from the Site's Own Bundle: 7 Requests over 533 Pages

Adityo Guni Waluyo

Finding a site's own GraphQL endpoint in its public JS bundle: 5330 rows in 7 requests, not 533 HTML pages.

TL;DR

Expecting a tedious HTML scraping job for Indonesia's procurement blacklist, the author found an unauthenticated GraphQL endpoint exposed by the site's XLSX export. A perPage of 800 meant the full 5,330-row dataset downloaded in just seven requests and about ten seconds. The daily snapshot plan was retired in favor of a simple full-dataset sync.

My plan for the DemandScope blacklist monitor sounded reasonable: take a daily snapshot of the first HTML page of the public procurement blacklist portal. In my head I was writing a script that would page through hundreds of listings, waiting for JavaScript to render each table before scraping it. Slow, and fragile by design. My assumption: the deeper data sat behind either traditional pagination or a tightly guarded API. I expected the standard defenses: a captcha, cookie auth, maybe CSRF tokens on every request. Reading the public JavaScript bundle of that Next.js client app said otherwise. None of those defenses existed. There was a named query, GetBlacklists, fired straight at /graphql [10]. And the best hint came from the site's own XLSX export button, which happily revealed the sanctioned request shape including perPage: 800. That fits how the framework works. Next.js renders layouts and pages as Server Components by default, but client-side apps lean on Client Components for state and browser-side data fetching [8]. With no auth, captcha, or cookies in the way, one request with 800 per page returned page 1 of 7: 5330 rows in total. The full sweep took 7 requests, roughly 10 seconds, with a bot User-Agent paced at one request per second. The whole dataset in one command instead of 533 separate HTML pages. And the dataset is far richer than the visible table: masked NPWP taxpayer numbers, provider addresses, pagu/hps contract values, violation regulation text, LPSE/KLDI/satker origin, SK numbers, publish and expiry timestamps. From my recount of the committed 7.7 MB JSON: 5330 rows, with 348 PUBLISHED, 4916 EXPIRED, 63 CANCELED, and 3 CANCELED_TEMPORARY. Masked NPWP numbers appear on 4169 rows. Tender pagu or hps values exist on 3006 rows, about 56 percent. Start dates span 1905, one bogus row, to 2026. 27 rows start in September-October 2026, exactly the fresh-signal window a daily digest needs. 4260 unique provider names in total. Offset pagination has known performance and security downsides on large datasets [9]. At 5330 rows this one is small, so seven honest requests beat any cleverness. GraphQL introspection would even map the schema's types and fields without guessing [11]. This discovery retired the daily page-1 snapshot plan entirely. Ingest is now a plain HTTPS GraphQL client, bot User-Agent, one request per second, syncing the full dataset daily in 7 requests and deduplicating by row id. The 5330-row dump is the day-one historical baseline; older rows never get re-fetched. The HTML parsing lane is retired, and the page-1 JSON survives only as lane provenance. The request pattern is short:
import requests

url = "https://daftar-hitam.example.id/graphql"
headers = {
    "User-Agent": "DataBot/1.0 (Research Purpose)",
    "Content-Type": "application/json"
}

payload = {
    "query": """
    query GetBlacklists($page: Int, $perPage: Int) {
        blacklists(page: $page, perPage: $perPage) {
            data {
                id
                providerName
                status
                npwp
                paguHps
                startDate
                endDate
            }
            pageInfo {
                currentPage
                totalPages
                totalData
            }
        }
    }
    """,
    "variables": {"page": 1, "perPage": 800}
}

response = requests.post(url, json=payload, headers=headers)
print(response.json())

Sources

  1. Next.js Docs: Server and Client Components
  2. GraphQL Learn: Pagination
  3. GraphQL Learn: Serving over HTTP
  4. GraphQL Learn: Introspection

Related articles