Closing the Data Leak in an Automated Article Pipeline
Briefs headed to an external chatbot still carried hostnames and real repo names. A three-checkpoint mechanical gateway shut it.
TL;DR
Raw internal data was leaking to an external LLM, and the fix wasn't better prompts but a mechanical architecture. The pipeline now sanitizes at generation via fictional alias mapping, blocks pushes if real tokens remain, and continuously scans published content. It uses IETF reserved placeholders like example domains and documentation IPs, proving sanitation is a supply-chain control.
Automated article pipelines often leak internal data to external LLMs. The solution is not prompt engineering, but architecture: a mechanical gateway that rewrites real identities into fictional brands before data leaves the system. **Slug (untuk sistem)**: gerbang-merek-fiktif-en
I was checking the pipeline logs that ship briefs to an external chatbot. Scrolling through the output, my eyes caught one line: production hostname, repository paths, and client names sitting right there in the text about to be processed. Initially, I guessed this was a prompt engineering issue. Maybe the prompt lacked strictness or the instructions were ambiguous. But after tracing the data flow, the problem was more fundamental. Our internal data was reaching the external model in its raw form. If prompt injection is an inbound threat, this is an outbound one — and I realized OWASP LLM01:2025 already covers this bidirectional risk [4].
Three mechanical checkpoints
Trusting a model's discipline is not a control. Models are probabilistic, not deterministic. So I built a mechanical gateway at three points. First, at the generate stage. Before the brief touches the external LLM, a rewriting step transforms real entities into fictional ones using an alias map. This map is injected first, so the output is already sanitized. The real brand never reaches the context window. Second, at the push stage. There is a hard gate that exits with code 1 if the content, title, excerpt, SEO metadata, or slug still carries a real token. This is not a dismissible warning. If the gate rejects it, the push fails entirely before the article reaches the API, except for meta-articles discussing the system itself, where a specific comment downgrades it to WARN. Third, at the verify stage. A corpus watchdog runs continuously, scanning published articles. It currently operates in WARN mode pending calibration, but it has a strict option. The self-test includes 10 asserts covering regex boundaries, ensuring hyphen, slug, and URL segments match, and alphanumeric continuation blocks handle compound half-matches.An official palette that already exists
There is no need to invent new conventions. The IETF has provided an official replacement palette long before LLMs existed. RFC 2606 from 1999 reserves example.com, .test, and .invalid specifically for documentation and testing to reduce conflict [1]. For IP addresses, RFC 5737 from 2010 provides three specific blocks for documentation: 192.0.2.0/24 (TEST-NET-1), 198.51.100.0/24 (TEST-NET-2), and 203.0.113.0/24 (TEST-NET-3) [2]. These addresses should not appear on the public internet. For private networks, RFC 1918 reserves 10.0.0.0/8, 172.16.0.0/12, and 192.168.0.0/16 [3]. Mapping real identities onto this official palette is all you need. Consistency beats creativity every time.Field results
A baseline check of 564 published articles found 7 still carrying real tokens. The watchdog currently runs in WARN mode as I write this, pending final calibration of the strict threshold. But this number highlights a clear truth: identity sanitation is a content supply-chain problem, an architecture property, not a prompt-craft problem. RFC 6973 names this outward consequence as 'secondary use', meaning the use of collected information for a purpose different from the one it was collected for, which may violate people's expectations [5]. What never gets sent never leaks.Sources
- [1] RFC 2606: Reserved Top Level DNS Names (rfc-editor.org)
- [2] RFC 5737: IPv4 Address Blocks Reserved for Documentation (rfc-editor.org)
- [3] RFC 1918: Address Allocation for Private Internets (datatracker.ietf.org)
- [4] OWASP LLM01:2025 Prompt Injection (genai.owasp.org)
- [5] RFC 6973: Privacy Considerations for Internet Protocols (datatracker.ietf.org)