Reading a Tech Stack from Public Signals
Passive recon: map a stack from subdomains, archived URLs, TXT records, MX, nameservers, and secret scans without sending one packet to the target.
TL;DR
Passive recon maps a domain's tech stack from public data alone—DNS, subdomains, old URLs, headers—without touching the target. The catch: signals mislead easily, since verification tokens don't prove actual usage and Cloudflare headers matter less than nameserver records. Found secrets? Stop at discovery—validation crosses into active territory, so report through official disclosure channels.
Ten at night, my screen shows exactly one domain: cakradata.example, a fictional company I use as the example throughout this article. Zero packets sent to anyone's servers. No port scans, no direct requests, no interaction with the target infrastructure. What I have is a handful of browser tabs full of public data and one terminal window. From that alone, I can piece together a credible picture of the tech stack behind the domain.
Halfway through, I nearly fell into one wrong conclusion. There was a TXT record in the DNS starting with anthropic-domain-verification-. My first guess: this company runs Claude in production. The reality is duller. The token only proves the domain owner once verified ownership of the domain with Anthropic for an SSO or provisioning flow. WorkOS is Anthropic's provider for domain verification, and it is the customer who is told to place that TXT record in their own DNS [1]. API keys used for inference never leave a trace in public DNS. Ownership verification is not proof of a working integration. That is the classic passive recon trap: a conclusion that skips one inference step too many.
Mapping the surface through third parties
The earliest step: subdomains. subfinder collects subdomains from passive online sources, from certificate transparency logs to security datasets, without sending a single packet to the target [3]. Its passive design is marketed as speed plus stealth [3], but the value I care about is simpler: my queries land on third-party servers, not on the target's front door.
The second layer: old URLs. gau fetches known URLs from AlienVault's Open Threat Exchange, the Wayback Machine, Common Crawl, and URLScan for a given domain [5]. This historical footprint often exposes forgotten staging endpoints or API route patterns that accidentally reveal the backend framework. OWASP's WSTG treats this whole family, from search-engine discovery to fingerprinting and architecture mapping, as standard, documented information gathering [6]. This is not dark-age trickery; it has a manual.
DNS: honest records, easy misreads
The first two commands I run against the example domain:
# vendor verification tokens and other TXT records
dig TXT cakradata.example +short
# inbound mail routing
dig MX cakradata.example +short
MX records give a relatively stable signal for the email vendor. The current Google Workspace value is smtp.google.com. But domains that started using Google Workspace before 2023 may still carry the older five-record set beginning with aspmx, and Google states those legacy values are still supported [2]. A fingerprint that only recognizes the new value misreads legacy domains as non-Google. Match both sets before concluding anything.
Headers lie, nameservers rarely
On the HTTP side, the Cf-Ray header identifies the data center that processed the request with a three-letter code [7]. That sounds like an unmistakable Cloudflare signature. It is not. On Shopify, whose own CDN is Cloudflare-backed, cf-ray or server: cloudflare headers appear on ordinary responses and prove nothing about a merchant's infrastructure choices [8]. The more reliable check lives in DNS delegation:
# who manages this domain's zone
dig +short NS cakradata.example
Cloudflare nameservers mean the zone is managed in a Cloudflare account, the prerequisite for their proxying service [8]. It is a strong signal, not absolute proof: partial CNAME setups can use Cloudflare without moving nameservers. So never conclude "no Cloudflare NS means no Cloudflare".
Leaked secrets and the line you do not cross
Public artifacts leak more than their owners intend. TruffleHog discovers, classifies, validates, and analyzes secrets across git, wikis, logs, object stores, and filesystems, with more than 800 detector types [4]. For a passive posture the rule is one sentence: stop at discovery. Validation means testing the credential against the vendor's API, and that is an active act that lands in someone's security logs. The moment you send an authentication request with a found secret, you are no longer observing.
On ethics and law: the DOJ's CFAA charging policy says prosecution should be declined when conduct amounts to good faith security research, meaning accessing a computer solely for good-faith testing, investigation, or correction of a security flaw, without harming anyone [9]. That is a charging policy in one jurisdiction, not a universal license. Honor your scope, never validate found secrets, and report findings through official channels. RFC 9116 standardizes security.txt so organizations can declare their vulnerability disclosure process and researchers can actually find it [10].
My honest opinion: public signals are cheap; interpretive discipline is expensive. The most dangerous failure in passive recon is not missing data, it is the conclusion that skips a step. And the distance between reading a footprint and touching someone's system is exactly one unnecessary request wide.
Sources
[1] https://support.claude.com/en/articles/13132885-set-up-single-sign-on-sso
[2] https://support.google.com/a/answer/174125
[3] https://github.com/projectdiscovery/subfinder
[4] https://github.com/trufflesecurity/trufflehog
[7] https://developers.cloudflare.com/fundamentals/reference/http-headers
[8] https://shopify.dev/docs/storefronts/themes/best-practices/performance/avoid-request-proxies
[9] https://www.justice.gov/opa/press-release/file/1507126/download