Resilient Outgoing Postback / Webhook Architecture — Design Spec
- Status: Draft for review
- Original Author: Ron White (ronco)
- Date: 2026-07-01
- Chosen approach: Option A — In-house, evolve (see Alternatives Considered)
- Companion ADR: Resilient Outgoing Postback Architecture (the decision record; this spec is its deep-dive). Supersedes ADR 0019.
1. Context & Problem
AdGem sends outgoing postbacks/webhooks to publishers when a conversion or progress goal is recorded. Over ~18 months this path has produced 11 Datadog incidents. Nearly every remediation was a bespoke one-off resend/backfill script, which incident timelines explicitly flagged as "an architectural issue."
Root causes (from Datadog; full detail in the companion ADR)
Known categories (all confirmed): publisher outage, malformed URLs, missing/renamed macro data (most common), V3 payload type changes, added V3 fields breaking strict publishers.
Categories found beyond the known list, each driving a design requirement:
- Queue payload / class-version skew on rolling deploys (IR-266, IR-272) — serialized typed DTOs
on the queue fail
unserialize()when producer and consumer run different code. - Our own edge/WAF blocking legit postbacks (IR-156).
- Upstream pipeline silently gating postbacks (IR-255) — conversions never reached "approved," so postbacks never fired.
- Source-IP / trust drift with the MMP (IR-126).
- Missing upstream mapping → permanent DLQ (IR-285) — retries can never succeed.
- New offerwall (Prism) bypassing the legacy click-URL pipeline (IR-283).
- Upstream data-integrity fault delivered as real money (IR-307) — duplicate
tune.offerrows from two Airbyte streams doubled theplayer_sub_goalsjoin beforeROW_NUMBER(), halving goal repetition thresholds. The Turbo DAG created 30,083 conversions (8,797 retrodated as far back as March 18); auto-approval keyed on load time rather than conversion datetime, so they approved. Nothing on the postback path was broken — it delivered ~$85,132 of $85,466 in payouts across 41 publishers exactly as designed. Detection was an employee noticing a dashboard, not a monitor.
Current architecture pain (api @ origin/master)
- One job (
SendOfferConvertedWebhookV2) → factory →V2Webhook(GET+macros) /V3Webhook(POST+JSON) /PaytronixWebhook. - Retry = 3 in-process attempts with blocking
usleep(ties up a worker up to ~65s); jobfailed()is a deliberate no-op → no DLQ, no queue-level backoff. - Failures land in
failed_webhooksand sit until a human runs the admin API / CLI. Replay re-renders from the stored URL/body with a live-fetched key → old-URL/new-key skew. - Silent V2 macro failures: an unknown/renamed macro becomes an empty string — no log, no metric, no failure. Ships a broken-but-2xx postback that looks successful.
- Two macro lists that can drift, plus a third copy client-side in the dashboard.
- Dual signing code paths behind DevCycle flag
use-refactored-webhook-services. This flag has been at 100% since Feb 2026 (confirmed via DevCycle); the inline branch is dead code. - No server-side URL/payload validation; the dashboard "Test Postback" is unwired on master.
- Adding a new hook type needs three edits, one cross-repo (enum in
adgem/common, hardcoded factory array, 5 abstract methods) — not "minimal effort."
Production volume (14 days, env:production — also in the ADR appendix)
| Metric | Avg/day | Peak day | Peak rate |
|---|---|---|---|
| Conversions | ~51.6K | 58.3K | ~1/sec |
| Outgoing sends | ~72K | 79.7K | ~1.5/sec sustained, ~30–50/sec burst |
| Send success | 99.95% | — | 443 exhausted / 14d |
| Retries | ~0.11% of sends | — | — |
| Distinct apps | ~150 | — | — |
Sizing conclusion: this is a low-volume workload. Everything stays in PostgreSQL; no DynamoDB
needed. postback_deliveries ≈ 80K rows/day peak (~29M/yr); delivery_attempts ≈ 150 rows/day.
Partitioning + retention is hygiene, not load-critical.
Fan-out note: sends exceed conversions ~1.4× because one conversion can produce multiple
postback-worthy events (progress goals + final reward), each hitting the app's single endpoint
(outgoing_postback_settings is one row per app). It is not multi-endpoint fan-out today.
2. Goals & Non-Goals
Goals
- Self-serve replay + edit (headline): view failed postbacks and re-trigger them without engineering — including editing a bad URL/macro/field so the retry actually succeeds. Internal support/ops first (Phase 1), scoped publisher-facing view later (Phase 2).
- Fix-once, re-render-many: always render from the immutable source + current config, so one config/code fix makes the whole backlog succeed on re-drive.
- Durable delivery: real DLQ + out-of-worker backoff; no blocking
usleep; no silently-dropped failures. - Blast-radius isolation: one bad publisher can't degrade delivery for everyone (per-pub circuit breaker + quarantine).
- Safe payload evolution: versioned V3 schema + per-publisher field subsetting, so additive changes are non-breaking by default.
- Double-reward prevention: a stable idempotency key across retries/replays/edits.
- Cheap extensibility: adding a custom partner hook (e.g. LINE) = one class + register.
- Observability of the silent gaps: catch "approved-but-not-delivered" and render errors, not just HTTP status codes.
- Economic containment: delivery is the last gate before the publisher is told, so an abnormal payout rate — or a retrodated source event — halts delivery for human review rather than shipping irreversibly (§5.11). Making that gate meaningful also requires payout-linked analytics to follow delivery rather than precede it (§5.11), which is a change to the current ordering.
Non-Goals
- Rewriting the conversion/approval pipeline. Root causes upstream of us stay upstream: IR-255's silent gating and IR-307's duplicate-source join both need fixing where they happen (for IR-307, source-level uniqueness validation in the Airbyte/Turbo path). What we add here is the last-gate half — we can detect the gap (§9) and now refuse to propagate an anomalous one (§5.11), but we cannot prevent a bad approval from being made.
- Full free-form publisher payload construction (deferred; see Alternatives).
- Replacing Redshift as the analytics audit trail.
- Moving to a managed webhook vendor or a greenfield service (see Alternatives).
3. Design Principles
- Separate immutable source from mutable delivery. Source-snapshot columns are write-once; all rendering derives from them + current config. This is what makes fix-once/re-render-many true.
- The queue carries primitive IDs, never serialized typed objects. Workers re-hydrate from
Postgres. Structurally kills deploy-skew
unserializedeaths (IR-266/272). - Fail loud, never silent. A missing/renamed macro or schema mismatch is a render error that stops the send and raises a metric — never a blank-but-2xx postback.
- One source of truth per concept. One macro registry (shared by render, validation, and the dashboard picker); one signing implementation.
- Resilience lives in the base, variation lives in providers. Retry/DLQ/idempotency/logging are inherited; a hook provider only supplies render/transport/signing differences.
4. Architecture Overview
Substrate: Laravel workers (reusing all PHP render/sign code) + SQS for durable retry/backoff and a real DLQ; a cache-backed per-publisher circuit breaker provides logical isolation without physical per-pub queues (deferred escalation if head-of-line blocking ever survives the breaker).
5. Components
5.1 Outbound Hook Provider Registry (extensibility — LINE)
Replaces the hardcoded factory array + shared WebhookType enum with a registry of self-registering
providers. Each provider implements a small interface; the resilient send loop (retry, DLQ,
idempotency, logging, failure capture) lives in a shared base and is inherited.
- Adding a hook = one class + one
register()line. No cross-repo enum edit;hook_typebecomes a validated registry key onpublisher_endpoints. - Paytronix is refactored into the reference custom-provider; LINE is the next provider and the acceptance test for "minimal dev effort."
- Providers declare their transport (method, headers, auth) and signing, so a custom partner overrides only what differs and inherits everything else.
5.2 Renderer (loud, single source of truth)
- One canonical macro registry consumed by render + validation + the dashboard picker (collapses today's three drifting lists).
- Unknown/renamed/empty-resolving macro → render error → delivery
status=render_failed, metric emitted, surfaced for attention before any HTTP call. - Renders from the delivery's write-once
source_snapshot+ the endpoint's current config.
5.3 Delivery worker + SQS DLQ
- Consumes primitive IDs, hydrates the delivery, renders, signs, sends with an HTTP timeout.
- Transient failures use SQS visibility-timeout backoff (no in-process
usleep); exhausted deliveries land in the DLQ and are surfaced in the replay UI — never silently dropped. - Retries write an attempt row; success is a status flip (no row) to keep
delivery_attemptssparse.
5.4 Per-publisher circuit breaker
- Cache-backed error-rate per app/publisher. When it trips, that pub's deliveries divert to a quarantine queue with delayed re-drive, so a downed/erroring pub neither hammers itself nor starves workers for everyone else.
5.5 Idempotency (double-reward prevention)
- A stable idempotency key derived deterministically from the delivery target and event
(
endpoint_id+conversion_id+event_type/goal_id+app_id). Scoping toendpoint_idkeeps deliveries distinct when the schema grows to multiple endpoints per app (today one endpoint per app, but the model allows more). Unlike today's per-attemptrequest_id, it does not change across retries, replays, or edits. - Unique on
postback_deliveries(UNIQUE(endpoint_id, conversion_id, event_type)) → internal double-dispatch (e.g. support re-fire racing the original) collapses to one delivery per target. - Sent to publishers as an opaque keyed hash —
X-AdGem-Idempotency-Key = HMAC(secret, key)— never the raw internal identifiers, so we can dedupe without leakingapp_id/conversion_idstructure. Stable in the payload too, so publishers can dedupe → prevents double-reward on their side when we re-drive a delivery that actually succeeded.
5.6 Validation & test-send (pre-flight)
- Server-side validation on endpoint save (dashboard → api): URL syntax/scheme, macro names
resolve against the registry, and for V3 a payload-schema dry-render against the pinned
schema_version. Bad config is rejected at config time. - Revived test-send: renders a sample event through the real render+sign path and shows the
publisher the exact bytes + the receiver's response. (Today's
sendPostbackRequestis unwired.) - URL validation includes an SSRF allow-list / deny-internal-ranges check (also applied to edits).
5.7 Response validation (per-publisher, post-send)
Today success is $response->successful() — a 2xx and nothing else
(OfferConvertedWebhook.php:84). The response body is already read, logged, sanitized into the
Redshift outgoing-postback stream and persisted to FailedWebhook.response_body — but nothing
parses it. IR-258 exposed that three publishers on the same code path disagree about what a
response means:
| Publisher | Success signal | Failure signal | Under status-only validation |
|---|---|---|---|
| Capital One | 2xx | proper 4xx/5xx | Correct. Strictest validator we have. |
| Dainata | 2xx | no proper error | Failures look like successes. |
| Partnerize | 2xx + conversion ID in body | 2xx + empty body | Failures look like successes. |
Partnerize has confirmed in writing they will not change this ("this is how the logic is currently built into our system, and we won't be able to modify this behaviour"). So "publishers should return correct status codes" is a position we can advocate, not a design we can depend on.
Design. A response_validator strategy resolved per endpoint, using the same registry pattern
as hook_type (§5.1): validated on save, exercised by test-send (§5.6), one implementation per
strategy.
| Strategy | Passes when | Covers |
|---|---|---|
status_only | 2xx (current behaviour) | Default. Capital One and every existing endpoint |
non_empty_body | 2xx and body non-blank after trim | Partnerize |
body_matches | 2xx and body matches a configured pattern | Partnerize, tighter than non-emptiness |
json_field | 2xx and a named JSON field exists / equals / differs from a configured value | {"status":"error"} returned with a 200 |
Config lives on publisher_endpoints as response_validator (registry key) plus
response_validator_config (JSON). Existing endpoints default to status_only, so this is
non-breaking for all ~150 apps currently receiving postbacks. A hook provider may also declare its
own default validator, so a Partnerize-shaped custom provider isn't left misconfigured by omission;
per-endpoint config overrides the provider default.
Retry semantics — the subtle part. A validation failure is not automatically safe to retry. During IR-258 Dainata declined a resend specifically because of double-reward danger: when a publisher gives us no reliable failure signal, we cannot distinguish "they never got it" from "they got it and didn't say so". Therefore:
- A validator failure resolves to its own state,
validation_failed, rather than falling into the transient-retry path. - Auto-retry is a per-endpoint decision (
on_validation_failure: retry | hold), defaulting tohold— surfaced in the replay UI for an idempotency-aware human decision. - The idempotency key (§5.5) is what makes
retrysafe where the publisher honours it;holdis the honest fallback where they don't.
Either way it is alertable: a per-endpoint validation_failed rate is precisely the signal that
was missing for the ~3 months IR-258 ran undetected behind a feature flag.
Non-goal. This cannot rescue a publisher whose success and failure responses are byte-identical. It makes the distinguishable cases detectable, and makes the indistinguishable ones an explicit, visible property of that endpoint's configuration rather than a silent assumption.
5.8 Payload schema versioning + per-pub field subsetting (the "middle path")
- V3 stays a flat, canonical, typed, versioned envelope → signing stays tractable; type changes
(IR-263) are handled by pinning each endpoint to a
schema_version. - Each endpoint opts into which canonical fields it receives (
field_subset, with optional key aliasing). New fields default to not-included for existing pubs, so additive changes are non-breaking without a version bump. Version bumps are reserved for type/structure changes. - Full free-form BYO payload construction is deferred (see Alternatives).
5.9 Suppression / poison handling
- Permanent failures (missing upstream mapping — IR-285) are classified and moved to
postback_suppressionsrather than retried forever; shown separately in the UI with suppress/resolve actions.
5.10 Replay + edit engine (the headline)
Shared engine, two audiences (see §8).
5.11 Payout-rate anomaly halt (economic circuit breaker)
Scope note. This section assumes the target analytics ordering described under Payout-linked analytics must follow delivery below. Under today's ordering the halt still prevents the postback, but not the records that assert it happened.
Distinct from §5.4. That breaker asks "is this publisher healthy?"; this one asks "do these payouts make sense?" — an error-rate breaker is blind to IR-307, where every send was a clean 2xx. Different signal, different scope, different clearing policy; they share only the divert-and-hold mechanism.
Trip signals (both configurable, both armed):
-
Payout value rate vs. trailing baseline. Payout value per rolling window compared against a trailing baseline for the same scope, seasonalized by time-of-day and day-of-week (postback volume is strongly diurnal — see §1's ~1.5/sec sustained vs. ~30–50/sec burst). Trips on a configured multiple over baseline, with an absolute floor so low-volume publishers don't trip on noise.
This signal already has a metric.
webhook.payout.usd(added toapiSept 2026) is a distribution carrying each outgoing postback's USD payout, taggedappId/webhookType/outcome, emitted once persend()rather than per attempt so retries don't scale the dollars, withsuccessandexhaustedseparating money that reached the publisher from money that didn't. The detect-only phase is therefore mostly baseline-building on an existing metric rather than new instrumentation.Its identity is not endpoint-resolvable, and the breaker is endpoint-scoped. The metric's identifying tag is
appId. That resolves to exactly one endpoint only while the current one-row-per-app invariant holds — and §6 deliberately does not preclude many endpoints per app, so this is a dependency on something the design intends to relax, not a stable property. Per-endpoint baselines built onappIdwould silently blend endpoints the day that changes. Requirement: before per-endpoint enforcement is armed,webhook.payout.usdmust carry an endpoint-identifying tag (or the one-endpoint-per-app constraint must be enforced in the schema rather than merely observed). Global-scope baselines are unaffected, and the retrodating guard is per-delivery so it does not depend on metric tagging at all — another reason it arms first. -
Source-age (retrodating) guard. A per-delivery shape check rather than a rate check: hold when
source_created_at − source_completed_atexceeds a configured window (both defined below). Stating the rule in named fields rather than as "completion precedes creation" is deliberate — IR-307 turned on auto-approval being keyed on load time instead of conversion datetime, so the rule that catches timestamp confusion must not itself be ambiguous. IR-307's cohort was defined by exactly this property (>7 days), so this rule catches the first delivery instead of waiting for an aggregate to build. Cheap, deterministic, and independent of baseline calibration — it is the guard that would have worked on day one.
Required snapshot inputs. Both signals depend on three values being resolvable from the delivery:
| Input | Meaning | Used by |
|---|---|---|
payout_value | Publisher payout amount for this event, in a single normalized currency | Value-rate signal |
source_created_at | When our pipeline created the conversion row | Retrodating guard |
source_completed_at | When the player actually completed the goal upstream | Retrodating guard |
Retrodation is source_created_at − source_completed_at. These three are a required subset of the
source_snapshot field set — the wider field set is still an open decision (§13), but §5.11 cannot be
built without these. Normalizing currency at snapshot time matters: a value-rate baseline computed
across mixed currencies is meaningless.
Missing-input behavior: fail closed — but only under enforce. If any required input is absent or
unparseable, the delivery goes to held_anomalous with a distinct missing_anomaly_inputs reason and
its own metric — never a silent pass. A guard that waves through what it cannot evaluate reproduces
exactly the failure mode IR-307 demonstrated.
This is scoped by mode (§6), and the distinction matters more than it looks:
mode | Missing required input |
|---|---|
off | Not evaluated. |
detect | Emit the missing_anomaly_inputs metric and deliver normally. Detect never diverts — that is what makes it safe to switch on. |
enforce | Divert to held_anomalous. |
Without that split, arming detect-only would hold every delivery whose snapshot predates these fields,
which is the opposite of a safe rollout. It also makes detect do real work: the missing-input rate it
reports is the readiness signal for arming enforce, and that rate should be near zero first. Expect
it to be high across historical failed_webhooks rows, whose snapshots predate these fields (Risk 5) —
those are released by review, not by weakening the rule.
Scope — evaluated at two levels simultaneously:
| Scope | Trips | Rationale |
|---|---|---|
| Per publisher endpoint | That endpoint's deliveries divert to held_anomalous | A single mispriced offer or runaway campaign concentrated on one pub |
| Global (all outgoing) | All outgoing delivery halts | IR-307's shape: $85.5K spread thin across 41 publishers, where no single per-pub threshold need ever trip |
The global scope is not redundant with the per-pub one. Broad-but-shallow is the failure mode a per-publisher breaker structurally cannot see.
Halt semantics. A trip does not drop anything. The postback_deliveries row and its write-once
source_snapshot are written as normal and the delivery moves to held_anomalous. The existing
replay engine (§5.10, §8) already knows how to re-drive from a snapshot, so a released cohort renders
from current config and ships unchanged — a halt costs latency, never data. held_anomalous is
semantically distinct from its neighbours: failed means we couldn't deliver, validation_failed
means the publisher signalled a problem, held_anomalous means we don't trust the economics of the
payload we're about to send.
Never publisher-visible. held_anomalous is excluded from the Phase 2 publisher-facing view (§8)
and from publisher-triggered replay, without exception. The state means we do not trust these payouts;
surfacing them to the publisher who would receive them hands release to the party with the strongest
interest in releasing. "Manual clear" means our human, and the exclusion belongs in the query that
backs the publisher view rather than in UI affordances.
Clearing — manual only. A reviewer with release authority works the held cohort in the replay UI
and either releases it (idempotency-guarded re-drive, triggered_by recorded) or rejects it
(resolved, never sent). There is deliberately no auto-resume: on a money-moving path a timer just
restarts the flood after the window elapses. The counterweight is that held volume and held age are
themselves alertable, so a forgotten breaker surfaces as a page rather than a silent revenue stall.
Reject does not reconcile itself. resolved, never sent closes the delivery but nothing upstream:
the conversion stays approved, and — under today's ordering — publisher-revenue is already emitted and
the reward already redeemed. There is no publisher-earnings ledger to reverse; settlement is publishers
invoicing us. So a reject must explicitly target the analytics record and the player-api reward, or the
discrepancy migrates onto the publisher's invoice and the player is left unrewarded with no postback to
explain it. Moving payout-linked analytics behind delivery shrinks this to the player-api reward alone;
until then a reject is a two-system operation and should be specified as one.
Payout-linked analytics must follow delivery. A hold is only as good as the records it holds back,
and today's ordering leaks around it. Verified against AdGem/api@origin/master
(IncomingPostbacksController), the payable path runs:
| Order | Step |
|---|---|
| 1 | publisher-revenue AdActionEvent — carries the publisher payout amount |
| 2 | playerApi->storeConversion(...) |
| 3 | LineConversionHook::issuePointsFor(...) |
| 4 | SendOfferConvertedWebhookV2::dispatch(...) — queued, fire-and-forget |
| 5 | playerApi->redeemReward(...) |
| 6 | reward-earned AdActionEvent |
Everything asserting the payout happened precedes the send, and step 4 is a queue dispatch whose
outcome nothing downstream waits on. So a held_anomalous delivery still leaves a publisher-revenue
record, a stored conversion, and a redeemed reward behind it.
Target ordering: events that assert a payout occurred — publisher-revenue, reward-earned —
are emitted by the delivery worker on confirmed delivery, not by the request handler on intent. Events
true regardless of whether the publisher was told — goal-complete, adgem-revenue — stay where they
are. This also fixes a second-order problem: the approved-but-not-delivered reconciliation in §9
compares Redshift against Postgres, and today Redshift is populated on intent, so it cannot distinguish
"approved and delivered" from "approved and merely attempted".
This is a change to api's current sequence, tracked separately from this spec. The ADR states the
target; the migration is its own work, and may be constrained by how entangled the current path is.
Relationship to existing controls. Two upstream anomaly controls already exist, both fraud-motivated: CAMP-184 (detection on sub-$25 conversions using Tune data, alerting only) and CAMP-120 (capping campaigns on abnormal conversion velocity). Neither caught IR-307: its conversions were high-value rather than low, and its volume was spread across 217 campaigns so no single campaign's velocity stood out. §5.11 operates at a different layer (delivery rather than Tune), on different signals (payout value and source age rather than volume), and is the only one of the three that can stop a postback. Complementary, not a replacement.
Rollout. Ship detect-only first — evaluate and emit metrics without diverting — and calibrate thresholds against replayed history (including the IR-307 window, where the correct answer is known) before arming the halt.
The two detectors do not enter detect-only together. The retrodating guard is per-delivery,
deterministic, and depends on neither currency nor metric tagging, so it calibrates immediately. The
value-rate detector has two unmet prerequisites, both open questions in §13: a defined currency for
payout_value and a unit contract for webhook.payout.usd, and an endpoint-identifying tag on that
metric. Calibrating value rate before the currency contract is settled produces a baseline that mixes
currencies — worse than no baseline, because it looks usable and would then be armed. So:
| Phase | Retrodating guard | Value rate |
|---|---|---|
| Now | detect — calibrate against replayed history | off — prerequisites unmet |
| Currency + unit contract settled | enforce | detect — begin calibrating |
| Baseline calibrated, endpoint tag present | enforce | enforce, per-pub then global |
Arming order therefore stays: retrodating guard, then per-pub value rate, then the global halt — but the gap between the first and second is a data contract, not just calibration time.
Honest limits. A trailing baseline cannot see an anomaly that ramps slowly beneath the threshold or
lands inside normal variance. The retrodating guard is only as good as the accuracy of
source_completed_at — note this is about a wrong timestamp, not a missing one: absence is decided
above (fail closed), while a plausible-but-wrong completion time defeats the guard silently and nothing
here catches that. And neither signal prevents a bad approval; they prevent an irreversible payout.
The cost of failing closed. Because the snapshot is raw pre-mapping under original field names, field presence varies by source, so a hook type whose upstream event simply has no completion timestamp would hold every delivery for that partner from the moment the guard is armed. That is the correct default — better a visible halt than a silently inert guard — but it must be caught at registration, not in production: a new provider added to the registry (§5.1) declares how it resolves the three required inputs, and the arming check refuses to enforce for a hook type that cannot supply them.
Attapoll, IR-307's most-affected publisher, avoided its loss by holding the conversions before approval. This is the same posture, one hop later and on our side of the wire.
6. Data Model
publisher_endpointsevolvesoutgoing_postback_settings(addshook_typeas a registry key,schema_version,field_subset, suppression + circuit state). One row per app today; the schema does not preclude many-per-app later.postback_deliveries— one row per dispatched postback-worthy event. Source-snapshot columns are write-once; status/attempt/circuit columns are mutable.delivery_attempts— sparse: retries, failures, and replays only.postback_suppressions— poison parking.payout_anomaly_settings— one row per (scope, detector), carrying both the thresholds and the arming state for §5.11. Rows exist for(global, value_rate),(global, retrodating), and per endpoint for each detector; the global rows are what make a system-wide threshold durable rather than an implicit constant. Config and state live together deliberately: an operator answering "why did this halt?" needs the threshold that was in force alongside the trip.modeis the arming switch, one per scope and detector, which is what lets the rollout in §5.11 arm the retrodating guard, then per-pub value rate, then the global halt independently.offskips evaluation;detectevaluates and emits metrics but never diverts (the detect-only phase);enforceevaluates and diverts toheld_anomalous. A missing row reads asoff, so a new endpoint is never silently enforcing an uncalibrated threshold.- Transitions.
modemoves only by deliberate config change (off → detect → enforce), audited; it is never changed automatically by a trip.circuit_statemovesclosed → trippedautomatically when a breach is seen andmode = enforce, andtripped → closedonly by a human (§5.11) — never on a timer and never on worker restart. Because clearing is manual, this state is durable in Postgres rather than cache-only; the cache in front of it is an optimization, and a cache flush must not re-open a halt.
- What
source_snapshotcaptures — snapshot the raw pre-mapping upstream event. Store the inbound event (e.g.TransactionData) under its original field names, before our DTO field mapping runs — not the post-mapping DTO. Rationale: a snapshot taken downstream of the mapping bakes in mapping bugs (IR-258'soffer_id/offerIdrename frozeoffer_id: null), so re-driving from it reproduces the bug. Snapshotting the raw event lets a mapping/rename fix be re-applied at render time, so a re-drive actually recovers. Trade-off: the renderer owns the field-mapping step (more render-path logic, slightly larger snapshots). See Open Questions for the field-set decision. - Data protection (governed by ADR 0044: Data Retention Policy).
source_snapshot,delivery_attempts.rendered_url/rendered_body/response_bodyhold PII (gaid/idfa/ip/state). Policy: store minimum-necessary fields; tokenize/redact sensitive values where the renderer can rehydrate from a system of record; encrypt at rest; and propagate deletions to hot tables and archives. Retention follows ADR 0044's 12-month PII window (anonymize/delete after). - Retention/partitioning: daily partitioning on the two high-volume tables, ~60–90 day hot operational retention for replay, older partitions archived to Redshift/S3 — all within the ADR 0044 ceiling.
- Future extension point: a parent
postback_eventstable (one per conversion) is introduced only if/when true multi-endpoint fan-out arrives (a likely trigger: the LINE work). - Redshift
outgoing-postbackAdActionEvents remain the long-term audit trail;delivery_attemptsis the operational/queryable store.
Delivery state machine
7. Idempotency & Double-Reward Prevention
See §5.5. Key invariants:
- The idempotency key is deterministic, scoped to the delivery target (
endpoint_id+ event), and stable across the delivery's entire lifetime (initial send, retries, replays, edits). - It is unique on
postback_deliveriesper target (internal dedupe) and transmitted to publishers as an opaque keyed hash — never raw internal identifiers (external dedupe). - Edits never mutate the source or the key — an edit produces a new
delivery_attemptwith a diff andtriggered_by, against the same immutable delivery.
8. Self-Serve Replay + Edit
Phase 1 — internal support/ops UI (replaces CLI + one-off scripts)
- Browse/filter undelivered deliveries by app, error_class, date, publisher.
- Inspect source snapshot, each attempt's rendered payload + response, and error classification.
- Two re-drive modes:
- Re-render & replay (default, safe) — re-renders from source + current config (fix-once path).
- Edit & replay (guarded) — override URL / macro mapping / specific fields for this re-drive; fully audited; optionally promoted to endpoint config so the fix sticks.
- Bulk ops with a circuit-breaker-aware drip (don't re-hammer a recovering pub).
- Poison deliveries shown separately with suppress/resolve.
Phase 2 — publisher-facing view (same engine, scoped)
- A publisher sees only their own undelivered deliveries, can fix their endpoint config + test-send,
and re-drive their own backlog. Edit scope is tighter (their config only), rate-limited, tenant-
isolated, and audited. Enabled by
triggered_byin the data model; deferred behind Phase 1. held_anomalousis excluded from this view and from publisher-triggered replay (§5.11). "Their own undelivered deliveries" would otherwise include the cohort we are holding because we don't trust its economics, letting the beneficiary release it. Enforce in the backing query, not the UI.
9. Observability
- "Approved-but-not-delivered" gap detector (highest value): reconciles approved conversions
(Redshift) against delivered
postback_deliveries(Postgres) on a batch cadence with a tolerance window. Catches IR-255 / IR-156 — failures where the postback code itself looked healthy. Scope limit: detects approved-but-not-sent, not should-have-been-approved. - Per-publisher success-rate SLO — feeds the circuit breaker and alerts before backlogs balloon.
- Render-error metric — makes the previously-silent macro failures loud.
- Payout-rate anomaly monitors — payout value per window vs. trailing seasonalized baseline, at
both per-publisher and global scope, plus a retrodated-source counter. These are the detection half
of §5.11 and run detect-only ahead of the halt. IR-307's detection method was
employee; this is the monitor that should have owned it. - Held-cohort depth and age — a manual-clear breaker fails dangerously if nobody clears it. Alert
on
held_anomalousvolume and on oldest-held age so a stalled halt pages instead of silently withholding legitimate payouts. Age is measured fromheld_at, set on entry to the state, not fromcreated_atorlast_attempted_at: a delivery can be created long before the breaker trips, or be held before it is ever attempted, so neither of those yields a correct hold age. - Circuit-breaker / DLQ / quarantine depth dashboards;
triggered_bybreakdown (auto vs support vs publisher re-drives).
10. Migration / Rollout
Incremental, parallel, reversible — no big-bang cutover.
- Delete the dead inline signing branch and retire
use-refactored-webhook-services(100% since Feb 2026 — this is cleanup, not a risky prerequisite). Standardize on the refactored services. - Additive schema first — create the new tables; start writing
postback_deliveriesin shadow on every postback-enabled conversion (no behavior change). - New render/sign path via the registry — dual-write attempts and compare bytes against the current path in shadow to prove signature parity before cutover. (Shadow = record+compare only; never send from both — that would double-postback.)
- Move retries to SQS DLQ, remove blocking
usleep, add the circuit breaker — per-app rollout via flag. - Support replay+edit UI ships against the new store; backfill existing
failed_webhooksintopostback_deliveriesso the current backlog is drivable. - Retire the old
failed_webhooksreplay path + duplicate macro lists once parity holds. - Phase 2 (publisher-facing view, test-send GA, field-subset UI) layers on after Phase 1 is stable.
11. Risks & Decisions (for the ADR)
- Idempotency / double-reward — resolved in design (§5.5, §7). Stable key across retries/replays/edits, unique internally, transmitted externally.
- PII in the source snapshot + retention/compliance — governed by ADR 0044: Data Retention Policy (12-month PII window, anonymize/delete after). Policy stance (above): minimum-necessary fields, tokenize/redact sensitive values, encrypt at rest, propagate deletions to hot tables and archives. Remaining decision: the exact snapshot field set (feeds the raw-pre-mapping snapshot choice) and the concrete delete-propagation mechanism — not whether we do these, only the specifics.
- SSRF / abuse on URL editing — edit-and-replay (and Phase 2 publisher self-serve) can change the outbound URL. Requires deny-internal-ranges allow-list, authz, tenant isolation, rate limiting, and full audit. Also applies to payload-field edits (data-exfil surface).
- Reconciliation feasibility — the approved-but-not-delivered detector crosses Redshift and Postgres. Decisions: batch cadence, tolerance window, query cost, and the explicit scope limit (can't catch should-have-been-approved).
- Backfill fidelity — historical
failed_webhooksrows may lack a full source snapshot, so some may only be replayable byte-for-byte, not re-render-from-source. Set expectations for the existing backlog vs. new deliveries. - Circuit-breaker escalation — if the logical per-pub breaker doesn't fully prevent head-of-line blocking at some future volume, escalate to physical per-pub queues (deferred).
- Payout-breaker calibration and authority — thresholds too tight halt legitimate payouts (revenue and publisher trust); too loose and IR-307 recurs. A trailing baseline is blind to a slow ramp, and the retrodating guard depends on a trustworthy completion timestamp in the snapshot. Mitigations: detect-only first, calibrate against replayed history including the IR-307 window, independently tunable per-pub and global thresholds, and alerting on held depth/age. Remaining decisions: who holds release authority for a held cohort, and the on-call response target for a global halt — a manual-clear control is only as good as the human process behind it.
12. Alternatives Considered (for the ADR)
The problem splits into rendering/signing/config (irreducibly AdGem's) and delivery mechanics (commodity — the only genuinely build-vs-buy part).
- Option A — In-house, evolve (CHOSEN). Grow
failed_webhooks+ admin API + Redshift audit into a real delivery subsystem. Keeps working code, lowest migration risk, no PII leaves our boundary, no per-message vendor cost, incremental, fits the phased audience plan. - Option B — In-house, greenfield (not chosen). New standalone delivery service, parallel-run, cut over. Cleanest end state but discards working signing/macro/audit code and adds cutover risk for little gain at this scale.
- Option C — Managed platform (Svix / Hookdeck / Convoy) (not chosen). Per-message pricing bites at our volume; PII would transit a third party (DPAs/residency); V3's byte-sensitive HMAC forces pre-signing and using the vendor as a dumb pipe, removing most of its value; the generic portal is not our "edit-the-macro-and-re-render" flow.
Deferred sub-alternative: full free-form publisher payload construction (BYO JSON). Rejected for v1 because arbitrary structure explodes the byte-sensitive HMAC + validation + support surface, and it doesn't solve type stability anyway. The chosen middle path (versioned schema + per-pub field subset + aliasing) captures most of the benefit safely.
Full comparison and evidence: the companion ADR (Considered Options + Appendix).
13. Open Questions
- Exact field set for the raw pre-mapping
source_snapshot(which inbound events/fields, and the tokenize/redact list) and the concrete delete-propagation mechanism (feeds Risk #2, governed by ADR 0044). - Reconciliation cadence/tolerance and where it runs (feeds Risk #4).
- Whether the LINE partnership introduces a second endpoint per app (would promote the
postback_eventsparent table from "future" to "now"). - Payout-breaker threshold values: the value-rate multiple and absolute floor per scope, and the
retrodating window (IR-307's cohort used >7 days; the right production value may be tighter). Where
they live is settled (
payout_anomaly_settings.thresholds); what they should be is not, and is to be set from replayed history during the detect-only phase (feeds Risk #7). - Currency normalization for
payout_value— which currency the value-rate baseline is computed in, and where conversion happens (snapshot time vs. evaluation time). Cross-currency baselines are meaningless, so this blocks the value-rate detector but not the retrodating guard. - Endpoint-identifying tag on
webhook.payout.usd— add one, or enforce one-endpoint-per-app in the schema. Blocks per-endpoint enforcement only; global scope and the retrodating guard are unaffected. - Release authority and on-call ownership for a held cohort — engineering, Publisher Management, or jointly (feeds Risk #7).
- Whether the payout breaker should also consult the upstream auto-approval confidence signal once the heuristic rule layer lands, rather than judging on payout shape alone.