Three Mapper Gaps Only Real Data Reveals
Green fixtures prove nothing: a nested object, camelCase keys, and full datetimes only surfaced once 5,330 real rows hit the pipeline.
TL;DR
A 5,330-row blacklist dump mapped error-free yet left company and date columns empty, because real rows broke three assumptions: nested provider objects, camelCase keys, and full RFC 3339 datetimes. Fixes layered a fallback resolver, one shared field registry, and regex date extraction. All 5,330 rows then ingested cleanly, confirming the lesson: start from row zero of real data.
The first rows of a 7.7 MB JSON dump carrying 5,330 blacklist entries had just arrived in the DemandScope ingest pipeline. Mapping ran without a single error, yet the company name column and the event date column came back empty. Inspecting row zero revealed three shape surprises at once.
The first guess blamed incomplete test fixtures. That guess was wrong in an uncomfortable way: every test was green precisely because the fixtures were built from the same assumptions as the mapper, so synthetic data could never expose a flawed assumption. The GraphQL schema fully describes the data a client may request [2], but assumptions about row shape live in client code, not in the schema.
Three Gaps in the Data Shape
The first gap was shape: the provider field was not a flat string but a nested object holding a name, a tax ID, and an address. The second was naming: date keys arrived as startDate and expiredDate rather than the snake_case style common in Python code. The third was precision: date values were full datetimes such as 2026-10-06T06:15:52Z, not bare dates.
All three make sense once the data provider's side is considered. The GraphQL naming guide recommends camelCase for fields to match JavaScript variable conventions [6]. What the portal sends are RFC 3339 datetimes, a profile of ISO 8601 [3]. Meanwhile, date.fromisoformat accepts strings in any valid ISO 8601 format with documented exceptions, and a full datetime with a zone marker falls under what it rejects [5]. That behavior is easy to verify directly:
from datetime import date
# Date-only format: succeeds
date.fromisoformat("2026-10-06")
# Basic format without separators: succeeds
date.fromisoformat("20261006")
# Full RFC 3339 datetime: raises ValueError
date.fromisoformat("2026-10-06T06:15:52Z")
Fixes and the Safety Net
The fix came in three layers. A resolver now reads the nested object first, then falls through to the existing flat-field chain, so the old lanes keep working. The camelCase key names were added to a single field-registry tuple, one source of truth consumed by both the parser and the detail-stripper; drift in that tuple had always meant dates disappearing silently, because two consumers would read different lists. Date extraction now runs a regex that lifts the date component before parsing, tolerant of space separators and optional seconds.
def resolve_provider_name(row: dict) -> str:
# Nested object first, then the flat-field chain
provider = row.get("provider")
if isinstance(provider, dict):
return provider.get("name", "").strip()
return str(provider).strip()
Post-fix verification ran on the same data: 5,330 of 5,330 rows were created with the normalized company column and the event date column filled [1]. An immediate re-ship produced 5,330 unchanged rows, revision counters untouched, only the last-seen timestamp moving; content-hash idempotency worked exactly as designed. Fifty-five tests stayed green. From 4,260 raw provider names, normalization left 4,141 unique companies, two numbers for two different layers.
The old RFC 1122 principle closes the lesson: be liberal in what you accept, conservative in what you send [4]. A mapper absorbs variance in shape, naming, and precision at the edge, then emits one canonical form downstream. One row carries a year of 1905 due to a typo in the source itself, and it ships as-is instead of being silently clamped; decisions about broken source data do not belong to the mapper. The Indonesian month-name lane showed the same pattern before, in the event_date column that was always NULL, and the lesson holds: start from row zero of real data, not from assumptions.