An Honest Contract for Batch Ingest Endpoints
One success number for a whole batch hides partial failure. Per-item created/updated/unchanged/rejected status makes it visible.
TL;DR
The new POST endpoint replaces a single batch status with per-item outcomes plus a summary so partial failures are visible. Guards cap payload size and item count, rate-limit per source, and fail closed with proper status codes including constant-time auth. Each item is marked created, updated, unchanged or rejected via a content hash, with an audit row for every request.
A single number hiding failure
My first attempt at the new endpoint played out exactly like the scenario I never wanted to inherit: the whole batch was declared done by a single status code, with no per-item story. Yet in an ingest pipeline, partial failure is normal — some items valid, some rejected over format. Returning one success status for the whole batch is a lie under those conditions: the sending system cannot tell total success apart from crippled success.
My initial assumption was that validating the entire payload before processing a single item could close the gap. For a continuous stream from external sources, that all-or-nothing approach just moves the problem: one bad item blocks hundreds of valid ones. The honest shape is independent per-item status, plus a batch-level summary.
This commit builds that contract at POST /v1/signals:ingest for DemandScope, guards and per-item semantics included, then tests it live: no token yields 401, a correct token creates new data, re-sending an identical payload counts as unchanged, and a changed payload is recorded as an update.
Custom-method route and fail-closed auth
The colon suffix on /v1/signals:ingest follows the custom-method convention: operations that do not map neatly onto standard CRUD verbs get their own room so the API vocabulary follows user intent [11]. A /v1/signals/ingest alias serves HTTP stacks that mangle colon paths. Both call the same handler.
Authentication uses the Authorization: Bearer header following the framework's documented convention [9]. The most-asked decision: why does a missing token configuration answer 503 instead of 401. It comes back to what the codes mean. 401 means the request lacks valid credentials [7], while 503 means the server is currently unable to handle the request due to a condition on its side, likely to ease after some delay [7]. A server whose token was never configured is a server-side condition, not the caller's mistake, so the endpoint fails closed with 503. The token comparison itself uses secrets.compare_digest, a constant-time comparison function built to reduce timing-attack risk [8].
Request guards: size, count, rate
Three fences guard this endpoint. First, a 1000000-byte cap enforced through a content-length pre-check; larger requests are refused with 413 Content Too Large, the code for content beyond what the server is willing or able to process [7]. Second, at most 100 items per batch. Third, a per-source rate limit of 60 requests per minute; overstepping answers 429 Too Many Requests, the code RFC 6585 introduced for clients sending too many requests in a given window [6], complete with a Retry-After header so the sender knows when to resume.
Every refusal above happens before a single item is processed. That is deliberate: cheap guards placed up front beat cleaning up half-ingested data.
Per-item semantics computed server-side
The core of the contract: every item earns one of four statuses, computed server-side from a canonical content hash keyed on the (source_id, source_ref) pair since migration 0003. The statuses are created for new items, updated for genuine change, unchanged when identical content is re-sent so only the last-seen time advances, and rejected for items that fail validation without dragging the rest down. Every request records one IngestRun row as an audit trail.
The response shape speaks for itself:
curl -X POST https://example.com/v1/signals:ingest \
-H "Authorization: Bearer " \
-H "Content-Type: application/json" \
-d '{"source_type": "sipp", "items": [{"dedupe_key": "item-001"}]}'
# Response:
# {
# "results": [{"source_ref": "item-001", "status": "created"}],
# "summary": {"seen": 1, "created": 1, "updated": 0, "unchanged": 0, "rejected": 0}
# }
With this structure, the sender can alarm only on rejected items, watch the unchanged ratio as a healthy-pipeline signal, and never guess what actually landed. Eleven new endpoint tests lock this behavior in, from 401 through the created-unchanged-updated cycle.
This commit left one working principle for any batch endpoint: never answer a batch with a single number. Per-item status granularity is not a luxury; it is what lets partial failure be seen, measured, and recovered automatically.