Authenticating...
Skip to main content

Outgoing Postback Macro Safeguards

Date: 2026-04-01 Original Author: Ron White Status: Draft

Related: The broader Resilient Outgoing Postback Architecture subsumes and generalizes these safeguards (loud render errors in place of silent empty macros). The tactical detection/resend runbook below remains useful in the interim.

Context​

On 2026-03-30, PR #3148 adopted the PlayerAPI path in OfferConverted. A property naming mismatch (offer_id vs offerId) in TransactionData caused offer_id to resolve to null for every postback sent through the new code path. PR #3188 fixed the naming on 2026-03-31.

Because publishers returned 200 even with empty offer_id values, we had no signal that anything was wrong. We need safeguards to catch this class of problem (any critical macro resolving to empty) both at send time and after the fact.

Goals​

  1. Detect missing critical macro values at the moment a postback is built, before it leaves our system.
  2. Monitor for anomalous changes in macro population across all apps, not just offer_id.
  3. Provide a runbook for identifying affected postbacks in Redshift and resending them.

Non-Goals​

  • Blocking postback delivery when a macro is empty. A degraded postback is better than a missing one.
  • Validating publisher-side data integrity. We control what we send; what they do with it is theirs.

Design​

Layer A: Build-Time Macro Validation​

Where: V2Webhook::build() and V3Webhook::build(), after the macro resolution loop.

What happens:

  1. After resolving all macros in the URL template, collect any macro that was present in the template but resolved to null or empty string.
  2. Check those against a critical_macros list in config/outgoing-postbacks.php. Initial set: offer_id, transaction_id, player_id, campaign_id.
  3. If any critical macro is empty:
    • Log a warning with app_id, campaign_id, conversion_id, and which macros were empty.
    • Increment a Datadog counter outgoing_postback.empty_critical_macro tagged with app_id and macro.
  4. Send the postback regardless.

Config addition (config/outgoing-postbacks.php):

'critical_macros' => ['offer_id', 'transaction_id', 'player_id', 'campaign_id'],

V2 implementation sketch:

After the existing macro replacement foreach loop in V2Webhook::build(), add:

$criticalMacros = config('outgoing-postbacks.critical_macros', []);
$emptyMacros = array_intersect($criticalMacros, $emptyResolvedMacros);

if (!empty($emptyMacros)) {
Log::warning('V2Webhook::build: critical macros resolved to empty', [
'app_id' => $offerConverted->app_id,
'campaign_id' => $offerConverted->campaign_id,
'conversion_id' => $offerConverted->conversion_id,
'empty_macros' => $emptyMacros,
]);

foreach ($emptyMacros as $macro) {
app(DatadogAdapter::class)->increment('outgoing_postback.empty_critical_macro', [
'app_id' => $offerConverted->app_id,
'macro' => $macro,
]);
}
}

The $emptyResolvedMacros array gets built inside the existing foreach loop by tracking macros where property_exists returns true but the value is null or empty string.

V3 follows the same pattern, but checks the JSON body fields rather than URL macros. In V3Webhook::build(), the offer_id and other fields are set directly from $offerConverted properties into the data array. The validation checks those properties for null before they're written to the body.

Layer B: Redshift Detection and Resend Runbook​

Detection approach:

  1. Query outgoing_postback_settings to find publishers whose webhook URL template contains {offer_id} (or any target macro). This tells us which publishers expect the value.
  2. Cross-reference against the outgoing-postbacks Redshift table. For V2 webhooks, look for resolved URLs where the macro value is empty (patterns like offer_id=& or offer_id= at end of URL). For V3, check post_data for "offer_id": null.
  3. Scope by time window: from when the flag was enabled to when the fix landed.

Sample Redshift query (V2):

SELECT
app_id,
campaign_id,
transaction_id,
conversion_id,
url,
created_at
FROM outgoing_postbacks
WHERE created_at BETWEEN '2026-03-30 00:00:00' AND '2026-03-31 18:00:00'
AND success = true
AND app_id IN (
SELECT app_id
FROM outgoing_postback_settings
WHERE url LIKE '%{offer_id}%'
)
AND (
url LIKE '%offer_id=&%'
OR url LIKE '%offer_id=$'
OR url REGEXP 'offer_id=[^a-zA-Z0-9]'
)
ORDER BY app_id, created_at;

Resend procedure:

  1. Run the detection query to get the list of affected conversion_ids and transaction_ids.
  2. For each transaction, fetch the correct offer_id from the PlayerAPI GET /api/v1/transactions/{txId}.
  3. Feed the corrected data through ResendWebhookService (app/Services/ResendWebhookService.php). The service already logs resends to the outgoing-postback AdActionEvent.
  4. Verify resend success by checking status codes in Redshift for the new attempts.

Datadog monitor:

Create a scheduled query monitor (or log-based metric if AdActionEvent data flows to Datadog Logs) that fires when outgoing postbacks contain empty critical macros. Group by app_id. Thresholds:

  • Warn: > 5 empty-macro postbacks in a 15-minute window for a single app.
  • Critical: > 1% of total postbacks for an app have empty critical macros over a 1-hour window.

Layer C: Per-App Macro Canary Metric​

Where: Same location as Layer A, inside V2Webhook::build() and V3Webhook::build().

What it emits:

For each macro that resolves to empty, increment a Datadog counter:

app(DatadogAdapter::class)->increment('outgoing_postback.empty_macro', [
'app_id' => $offerConverted->app_id,
'macro' => $macroName,
]);

This covers all macros, not just the critical set from Layer A.

Datadog anomaly monitor:

Set up an anomaly detection monitor on outgoing_postback.empty_macro grouped by app_id and macro. Use Datadog's built-in anomaly algorithm with:

  • Algorithm: Agile (responds quickly to sudden changes).
  • Deviations: 3 (avoids noise from minor fluctuations).
  • Minimum data window: 24 hours (prevents false positives for new apps that don't have a baseline yet).

The monitor fires when a macro that historically resolves to a value suddenly starts showing up as empty. This catches regressions for any macro, not just the ones we've been burned by.

Relationship to Layer A:

Layers A and C both emit metrics from the same code location, but serve different purposes:

  • Layer A's empty_critical_macro metric has a static threshold monitor. If offer_id is empty, alert immediately.
  • Layer C's empty_macro metric has an anomaly monitor. If any macro's empty rate deviates from its historical baseline, alert.

Layer A catches known-important macros. Layer C catches unknown regressions.

Testing Strategy​

  • Unit tests: Verify that V2Webhook::build() and V3Webhook::build() log warnings and emit metrics when critical macros resolve to empty.
  • Unit tests: Verify that the canary metric fires for any empty macro, not just critical ones.
  • Integration tests: End-to-end test that mocks the PlayerAPI response with a null offerId, flows through OfferConverted::createFromRewardViaPlayerApi, and asserts the warning log and Datadog metric are emitted.
  • Redshift query validation: Run the detection query against staging data with known-empty macro postbacks to verify it returns the expected rows.

Rollout​

  1. Ship Layer A (build-time validation) first. Low risk, immediate value.
  2. Set up the Layer B Datadog monitor and document the resend runbook.
  3. Ship Layer C (canary metric) and configure the anomaly monitor.
  4. Run the Layer B detection query against the 2026-03-30 to 2026-03-31 incident window and execute the resend runbook for affected publishers.