0076: Resilient Outgoing Postback Architecture
STATUS
Accepted
Supersedes 0019: Handle AdGem Outgoing Postbacks.
CONTEXT
AdGem sends outgoing postbacks/webhooks to publishers when a conversion or progress goal is recorded. Over ~18 months this path produced 11 Datadog incidents, and nearly every remediation was a bespoke one-off resend/backfill script — a pattern incident timelines explicitly flagged as an architectural problem. We want a system that is resilient by default and self-serve to recover from: failed postbacks should be viewable and re-triggerable without engineering, including editing a bad URL/macro/field so the retry actually succeeds instead of failing identically.
Deep-dive design: Resilient Outgoing Postback Architecture — Design Spec (data model, state machine, sequence diagrams, component detail).
Incident-driven requirements
Root-cause categories observed, and the design element each demands:
| Root cause (example incident) | Design response |
|---|---|
| Missing/renamed macro data (IR-258, IR-283) — most common | Render from source at replay; loud macro failures (never silent empty-string) |
| Queue class-version skew on rolling deploys (IR-266, IR-272) | Enqueue primitive IDs, hydrate in worker — no serialized typed DTOs on the wire |
| Upstream pipeline silently gating postbacks (IR-255) | "Approved-but-not-delivered" reconciliation monitor |
| Our own edge/WAF blocking postbacks (IR-156) | Same reconciliation + per-pub success SLO |
| Missing upstream mapping → permanent DLQ (IR-285) | Poison-message classification + suppression path |
| V3 payload type change (IR-263) / added fields break pubs | Per-pub pinned schema version + field subsetting |
| Publisher outage (IR-240) | Durable SQS retry + DLQ + per-pub circuit breaker |
| No self-serve recovery; bespoke scripts every incident | Self-serve replay + edit engine (support first, publishers later) |
Double-reward risk on resend (rotating request_id) | Stable idempotency key across retries/replays/edits |
| Adding a hook type needs 3 edits incl. cross-repo enum | Outbound Hook Provider registry (one class + register) |
| Publishers disagree on what a response means — failure arrives as an empty 200 (IR-258; Partnerize, Dainata) | Per-publisher response validation strategy (body and status) |
| Upstream produced valid-looking but economically wrong conversions, and we delivered them faithfully (IR-307) — $85.5K across 41 publishers, ~99.6% delivered before a human noticed | Payout-rate anomaly halt: value-rate breaker (global + per-pub) + retrodating guard, manual clear |
Current architecture (api @ origin/master), in brief
One job (SendOfferConvertedWebhookV2) → factory → V2Webhook (GET+macros) / V3Webhook
(POST+JSON) / PaytronixWebhook. Retry is 3 in-process attempts with blocking usleep; the job's
failed() is a deliberate no-op, so there is no DLQ and no queue-level backoff. Failures land in
failed_webhooks and wait for a human to run the admin API / CLI (see ADR 0056). Conversions do not
write to Postgres today (they are Redshift events); the only Postgres write on the conversion path is
FailedWebhook::create() on failure. The use-refactored-webhook-services DevCycle flag has been at
100% since Feb 2026, so its inline code branch is dead.
Considered Options
The problem splits into rendering/signing/config (irreducibly AdGem's — macros, v2/v3 payloads, byte-sensitive HMAC, per-pub config, PII, Redshift audit) and delivery mechanics (commodity — durable retry, DLQ, backoff, circuit-breaking, a failures portal). Only the second layer is genuinely build-vs-buy.
- Option A — In-house, evolve (DECISION). Grow today's
failed_webhooks+ admin API + Redshift audit into a real delivery subsystem, reusing all existing render/sign code. - Option B — In-house, greenfield. New standalone delivery service, parallel-run, cut over. Cleanest end state but discards working code and adds cutover risk for little gain at this scale (~80K sends/day peak).
- Option C — Managed platform (Svix / Hookdeck / Convoy). Per-message pricing bites at our volume; PII (gaid/idfa/ip/state) would transit a third party (DPAs/residency); V3's byte-sensitive HMAC forces pre-signing and using the vendor as a dumb pipe, removing most of its value; the generic portal is not our "edit-the-macro-and-re-render" flow.
- Deferred sub-option — full free-form publisher payload (BYO JSON). Rejected for v1: arbitrary structure explodes the HMAC + validation + support surface and doesn't solve type stability. The chosen middle path (versioned schema + per-pub field subset + aliasing) captures the benefit safely.
DECISION
Adopt Option A — In-house, evolve. The "buy" value (portal, DLQ) is mostly reachable by evolving what already exists; PII, cost, and byte-signing constraints make a managed platform awkward; and greenfield discards working code for little gain. Option A also matches the phased audience plan (one core serves internal support first, a scoped publisher view later).
Core design principle: separate the immutable source event from the delivery, and always render fresh from source + current config — which gives fix-once / re-render-many.
Key elements (detailed in the spec):
-
Substrate: Laravel workers (reuse PHP render/sign) + SQS for durable retry/backoff and a real DLQ; a cache-backed per-publisher circuit breaker for logical blast-radius isolation (physical per-pub queues deferred).
-
Data model (PostgreSQL):
postback_deliveries(one per postback-worthy event; write-once source snapshot + mutable status), sparsedelivery_attempts,publisher_endpoints(evolvesoutgoing_postback_settings:hook_type,schema_version,field_subset, suppression + circuit state), andpostback_suppressions. A parentpostback_eventstable is the extension point for future multi-endpoint fan-out. Thesource_snapshotcaptures the raw pre-mapping upstream event (original field names, before our DTO mapping) so mapping/rename bugs (IR-258) are recoverable by re-drive rather than frozen into the snapshot. -
Outbound Hook Provider registry: replaces the hardcoded factory + shared enum; a new custom hook (e.g. the LINE partnership) is one class + a
register()call, inheriting retry/DLQ/ idempotency/logging. Paytronix becomes the reference custom provider. -
Idempotency: a stable key scoped to the delivery target (
endpoint_id+ event), unique internally per target, and transmitted to publishers as an opaque keyed hash (X-AdGem-Idempotency-Key, never raw internal IDs), stable across retries/replays/edits → prevents double-reward and stays correct if an app gains multiple endpoints. -
Rendering: one canonical macro registry (shared by render, validation, dashboard picker); loud render errors; single signing implementation (delete the dead inline branch).
-
Validation & test-send: server-side URL + payload-schema validation on save; a revived test-send through the real render+sign path; SSRF allow-list on URLs and edits.
-
Payload evolution: versioned V3 schema + per-pub field subsetting; new fields default to not-included, so additive changes are non-breaking without a version bump.
-
Self-serve replay + edit: Phase 1 internal support/ops UI (re-render & replay; guarded edit & replay with audit and optional promote-to-config); Phase 2 scoped publisher-facing view.
-
Observability: approved-but-not-delivered reconciliation, per-pub success SLO, render-error metric, DLQ/quarantine dashboards.
-
Per-publisher response validation: success stops being "we got a 2xx". Each endpoint resolves a
response_validatorfrom a registry (status_only— the default, so nothing changes for existing publishers — plusnon_empty_body,body_matches,json_field), because publishers genuinely disagree about what a response means: Capital One returns proper HTTP errors, Dainata returns none, and Partnerize signals failure as an empty 200. A validator failure gets its ownvalidation_failedstate and is alertable; whether it auto-retries is per-endpoint and defaults to hold, since a publisher that can't signal failure also can't be safely re-sent to (Dainata declined an IR-258 resend precisely over double-reward risk). -
Payout-rate anomaly halt (economic circuit breaker). Every prior element assumes the event is correct and only delivery can go wrong. IR-307 broke that assumption: duplicate
tune.offerrows halved goal repetition thresholds upstream, the Turbo DAG created 30,083 conversions (8,797 retrodated as far back as March 18), auto-approval passed them, and this path delivered ~$85,132 of $85,466 in publisher payouts across 41 publishers before an employee — not a monitor — noticed. Delivery is the last gate before the publisher is told, so it gets a second breaker that trips on economics rather than errors. It is deliberately not the last gate before money moves — see the ordering requirement below, which this ADR changes:- Payout value rate vs. trailing baseline — payout value per rolling window against a time-of-day/day-of-week baseline, evaluated at two scopes at once: per publisher endpoint, and globally across all outgoing delivery. The global scope is what catches IR-307's shape; spread over 41 publishers, no single per-pub threshold need ever trip.
- Source-age (retrodating) guard — a per-delivery shape check, not a rate check: hold when
source_created_at − source_completed_atexceeds a configured window, wheresource_completed_atis when the player actually completed the goal upstream andsource_created_atis when our pipeline created the conversion row. Naming both explicitly is deliberate: IR-307 happened because auto-approval was keyed on load time rather than conversion datetime, and a rule whose job is catching that confusion must not itself be readable two ways. This is IR-307's exact signature and would have caught the first delivery rather than the aggregate.
Both signals depend on three inputs being resolvable from the delivery: payout value, source creation time, and source completion time. These become a required subset of the
source_snapshotfield set (which is otherwise still an open decision), and when one is missing the delivery fails closed intoheld_anomalouswith its own metric. A guard that silently passes when it cannot evaluate is precisely the failure IR-307 already demonstrated. Fail-closed applies underenforceonly: in detect mode a missing input is recorded and the delivery ships normally, since detect never diverts — otherwise switching on detection would itself halt every delivery whose snapshot predates these fields.Tripping diverts to a new
held_anomalousstate. The delivery row and its write-once source snapshot are still written, so the replay engine (element 8) re-drives the held cohort unchanged once released — a halt loses nothing but time. Clearing is manual only: a human releases the cohort (idempotency-guarded re-drive) or rejects it (resolved, never sent). No auto-resume — on a money-moving path a timer just restarts the flood. The held state is itself alertable so a forgotten breaker can't silently stall legitimate payouts.held_anomalousis never publisher-visible or publisher-releasable. It is excluded from the Phase 2 publisher-facing view (element 8) and from publisher-triggered replay. The state exists because we do not trust these payouts; exposing them to the publisher who benefits from them would let the party with the strongest interest in release perform it. "Manual clear" means our human.Rejecting leaves a divergence that has to be closed deliberately. Marking a cohort
resolved, never sentdoes not undo anything upstream: the conversion stays approved, thepublisher-revenueevent has already been emitted, and the player's reward has already been redeemed in player-api. Because there is no publisher-earnings ledger to reverse — settlement is publishers invoicing us — a reject has to target those two records specifically, or the discrepancy simply moves onto the publisher's invoice and the player stays unrewarded with no postback to explain it. The ordering change above shrinks this problem but does not remove it for anything already emitted.Payout-linked analytics must follow delivery, not precede it. Today they do the opposite, and that ordering is what makes a hold leaky. On the payable path in
api(IncomingPostbacksController,origin/master): thepublisher-revenueevent — carrying the publisher payout amount — is emitted roughly ninety lines before the postback is queued;SendOfferConvertedWebhookV2::dispatch()then queues the send fire-and-forget; andplayerApi->redeemReward()runs on the very next line, independent of the outcome. A held delivery therefore leaves behind an analytics record asserting the payout happened, a stored conversion, and a redeemed reward, with no postback behind any of them. Under this architecture, events that assert a payout occurred —publisher-revenue,reward-earned— move behind delivery confirmation, emitted by the delivery worker on success rather than by the request handler on intent. Events that are true regardless of whether we told the publisher —goal-complete,adgem-revenue— stay where they are. This is a real change to the current sequence and is tracked separately; the ADR states the target, not today's behaviour.Existing signal. The value-rate detector does not need a new metric:
webhook.payout.usd(added Sept 2026) is a distribution carrying each outgoing postback's USD payout, taggedappId/webhookType/outcome, emitted once persend()so retries don't scale the dollars, withsuccessandexhaustedseparating money that reached the publisher from money that didn't. The detect-only phase is largely a matter of building baselines on a metric that already exists. One gap: its identifying tag isappId, which resolves to a single endpoint only while the current one-endpoint-per-app invariant holds, and the data model deliberately allows that to change. The metric must carry an endpoint-identifying tag before per-endpoint enforcement is armed; global baselines and the per-delivery retrodating guard are unaffected.This is containment, not prevention: it does not stop bad approvals (IR-307's root cause is upstream, in the Airbyte/Turbo path, and its source-level uniqueness fix belongs there). It stops us propagating them irreversibly. Attapoll, the most-affected publisher, avoided the loss precisely by holding the conversions before approval — the same posture, one hop later.
Not our first anomaly control, and deliberately a different one. Two upstream controls already exist, both fraud-motivated: CAMP-184 (anomaly detection on sub-$25 conversions, on Tune data, alerting only) and CAMP-120 (capping campaigns on abnormal conversion velocity). Neither caught IR-307 — its conversions were high-value rather than low, and the volume was spread across 217 campaigns, so no single campaign's velocity stood out. §5.11 sits at a different layer (delivery, not Tune), on different signals (payout value and source age, not volume), and is the only one of the three that can stop a postback. They are complementary; this does not replace them.
CONSEQUENCES
Positive
- Recovery becomes self-serve and fix-once: fix config/code once, re-drive the backlog, everything renders correctly — replacing bespoke per-incident scripts.
- Structurally eliminates the deploy-skew failure class (primitive-ID queue payloads).
- One bad publisher no longer degrades delivery for everyone (circuit breaker + DLQ).
- Additive payload changes stop breaking strict publishers.
- New partner hooks (LINE and beyond) are cheap and inherit all resilience machinery.
- No silent delivery failures: render errors and approved-but-not-delivered gaps are alertable.
Note the limit on the response side — this holds only for endpoints configured with a body-aware
validator.
status_onlyremains the default, and astatus_onlyendpoint still reads any 2xx as success, so a publisher that signals a failed operation inside a 200 stays invisible until someone configures a suitable response-validation contract for it. - Publisher-signalled failures stop masquerading as successes, which is the gap that let IR-258 run undetected for ~3 months behind a feature flag.
- An upstream data-integrity fault that trips a configured signal stops converting straight into irreversible publisher payouts: for those, the breaker bounds exposure to one detection window instead of a full DAG run, and held deliveries stay replayable rather than lost. This is not a general guarantee — a fault that ramps slowly beneath the threshold or sits inside normal variance evades both signals and is bounded by nothing here (Risk 7).
Negative / costs
- New per-conversion Postgres write for postback-enabled apps (bounded: ~80K rows/day peak, sparse attempts, partition + retention). Low risk at current volume but real.
- We own the delivery mechanics we'd otherwise buy; requires disciplined SQS/queue-topology work.
- A migration period with shadow dual-write and byte-level signature-parity checks before cutover.
- New PII-bearing store (
source_snapshot) to govern (retention, encryption, deletion propagation). - The payout breaker can halt legitimate payouts. Manual-only clearing means a real trip needs a human on the other end, which is an on-call and runbook commitment, not just code.
Risks
- PII in
source_snapshot(gaid/idfa/ip/state) — governed by ADR 0044: Data Retention Policy (12-month PII window). Stance: minimum-necessary fields, tokenize/redact sensitive values, encrypt at rest, propagate deletions to hot tables and archives. Remaining decision: exact snapshot field set + concrete delete-propagation mechanism. - SSRF / abuse on URL & payload editing — deny-internal-ranges allow-list, authz, tenant isolation, rate limiting, full audit (especially Phase 2 publisher self-serve).
- Reconciliation feasibility — crosses Redshift and Postgres; define cadence, tolerance window, and the scope limit (catches approved-but-not-sent, not should-have-been-approved).
- Backfill fidelity — historical
failed_webhooksrows may lack a full source snapshot and be replayable only byte-for-byte, not re-render-from-source. - Circuit-breaker sufficiency — if the logical breaker doesn't prevent head-of-line blocking at higher volume, escalate to physical per-pub queues.
- Idempotency adoption — full double-reward protection on the publisher side depends on publishers honoring the idempotency key; internal dedupe holds regardless.
- Payout-breaker calibration — a trailing baseline is blind to an anomaly that ramps slowly
underneath the threshold or lands inside normal variance, and the retrodating guard is only as good
as the completion timestamp in the source snapshot. Thresholds too tight halt legitimate payouts
(revenue and publisher trust); too loose and IR-307 recurs. Mitigations: run detect-only first and
calibrate against replayed history before arming the halt, keep per-pub and global thresholds
independently tunable, and alert on held volume so a stalled breaker is loud. The two detectors do
not enter detect-only together: the retrodating guard calibrates now, while the value-rate detector
waits on a currency/unit contract for
payout_valueand an endpoint tag onwebhook.payout.usd— calibrating it sooner yields a baseline that mixes currencies, which is worse than none because it looks usable. Remaining decision: who holds release authority, and the on-call response target for a global trip. - Response validation has a floor — it can't help where a publisher's success and failure responses
are byte-identical, and a mis-specified validator could mark healthy traffic failed. Mitigations:
status_onlystays the default, validators are exercised by test-send before save, andon_validation_failuredefaults toholdso a bad validator creates visible held deliveries rather than a retry storm or double rewards. Per-endpoint config is also new operational surface for support to get wrong.
NOTES
References
- Supersedes: 0019: Handle AdGem Outgoing Postbacks
- Related / subsumed: 0056: Batch Resend Failed Webhooks (the new replay engine generalizes manual batch resend)
- Related: 0009: Postback Processor Phase 2 Architecture (incoming side; the SQS+DLQ pattern this ADR mirrors for outgoing)
- Deep-dive spec: Resilient Outgoing Postback Architecture — Design Spec
- Prior art: Outgoing Postback Macro Safeguards — build-time macro validation + anomaly canary for the IR-258 silent-macro class; subsumed and generalized by this ADR's loud-render-error design.
- Incidents: IR-126, IR-156, IR-219, IR-240, IR-255, IR-258, IR-263, IR-266, IR-272, IR-283, IR-285, IR-307 (payout spike — the driver for the payout-rate anomaly halt)
- Linear #169 — "design per-app postback response validation (handle body errors as well as HTTP errors)", the IR-258 postmortem follow-up (2026-08-04) satisfied by element 10 and spec §5.7. Design only; implementation is not committed by this ADR.
- PR #169: docs(adr): resilient outgoing postback architecture
- PR #219: docs(adr): halt postbacks on payout-rate anomaly
Original Author(s)
- Ron White (ronco)
Approval date
Approved by
Appendix
Production volume (Datadog, 14 days, env:production)
| Metric | Avg/day | Peak day | Peak rate |
|---|---|---|---|
Conversions (adgem_offerwall.completed_offer) | ~51.6K | 58.3K | ~1/sec |
Outgoing sends (webhook.dispatched) | ~72K | 79.7K | ~1.5/sec sustained, ~30–50/sec burst |
Send success (webhook.send.success) | 99.95% | — | 443 exhausted / 14d |
Retries (webhook.attempt) | ~0.11% of sends | — | attempt2=673, attempt3=456 / 14d |
| Distinct apps receiving postbacks | ~150 | — | — |
Sizing conclusion: low-volume workload; everything stays in PostgreSQL (no DynamoDB). Sends exceed conversions ~1.4× because one conversion yields multiple postback-worthy events (progress goals + final reward) to the app's single endpoint — not multi-endpoint fan-out.