Outgoing Postback Macro Safeguards
Date: 2026-04-01 Original Author: Ron White Status: Draft
Related: The broader Resilient Outgoing Postback Architecture subsumes and generalizes these safeguards (loud render errors in place of silent empty macros). The tactical detection/resend runbook below remains useful in the interim.
Context
On 2026-03-30, PR #3148 adopted the PlayerAPI path in OfferConverted. A property naming mismatch (offer_id vs offerId) in TransactionData caused offer_id to resolve to null for every postback sent through the new code path. PR #3188 fixed the naming on 2026-03-31.
Because publishers returned 200 even with empty offer_id values, we had no signal that anything was wrong. We need safeguards to catch this class of problem (any critical macro resolving to empty) both at send time and after the fact.
Goals
- Detect missing critical macro values at the moment a postback is built, before it leaves our system.
- Monitor for anomalous changes in macro population across all apps, not just
offer_id. - Provide a runbook for identifying affected postbacks in Redshift and resending them.
Non-Goals
- Blocking postback delivery when a macro is empty. A degraded postback is better than a missing one.
- Validating publisher-side data integrity. We control what we send; what they do with it is theirs.
Design
Layer A: Build-Time Macro Validation
Where: V2Webhook::build() and V3Webhook::build(), after the macro resolution loop.
What happens:
- After resolving all macros in the URL template, collect any macro that was present in the template but resolved to null or empty string.
- Check those against a
critical_macroslist inconfig/outgoing-postbacks.php. Initial set:offer_id,transaction_id,player_id,campaign_id. - If any critical macro is empty:
- Log a warning with
app_id,campaign_id,conversion_id, and which macros were empty. - Increment a Datadog counter
outgoing_postback.empty_critical_macrotagged withapp_idandmacro.
- Log a warning with
- Send the postback regardless.
Config addition (config/outgoing-postbacks.php):
'critical_macros' => ['offer_id', 'transaction_id', 'player_id', 'campaign_id'],
V2 implementation sketch:
After the existing macro replacement foreach loop in V2Webhook::build(), add:
$criticalMacros = config('outgoing-postbacks.critical_macros', []);
$emptyMacros = array_intersect($criticalMacros, $emptyResolvedMacros);
if (!empty($emptyMacros)) {
Log::warning('V2Webhook::build: critical macros resolved to empty', [
'app_id' => $offerConverted->app_id,
'campaign_id' => $offerConverted->campaign_id,
'conversion_id' => $offerConverted->conversion_id,
'empty_macros' => $emptyMacros,
]);
foreach ($emptyMacros as $macro) {
app(DatadogAdapter::class)->increment('outgoing_postback.empty_critical_macro', [
'app_id' => $offerConverted->app_id,
'macro' => $macro,
]);
}
}
The $emptyResolvedMacros array gets built inside the existing foreach loop by tracking macros where property_exists returns true but the value is null or empty string.
V3 follows the same pattern, but checks the JSON body fields rather than URL macros. In V3Webhook::build(), the offer_id and other fields are set directly from $offerConverted properties into the data array. The validation checks those properties for null before they're written to the body.
Layer B: Redshift Detection and Resend Runbook
Detection approach:
- Query
outgoing_postback_settingsto find publishers whose webhook URL template contains{offer_id}(or any target macro). This tells us which publishers expect the value. - Cross-reference against the
outgoing-postbacksRedshift table. For V2 webhooks, look for resolved URLs where the macro value is empty (patterns likeoffer_id=&oroffer_id=at end of URL). For V3, checkpost_datafor"offer_id": null. - Scope by time window: from when the flag was enabled to when the fix landed.
Sample Redshift query (V2):
SELECT
app_id,
campaign_id,
transaction_id,
conversion_id,
url,
created_at
FROM outgoing_postbacks
WHERE created_at BETWEEN '2026-03-30 00:00:00' AND '2026-03-31 18:00:00'
AND success = true
AND app_id IN (
SELECT app_id
FROM outgoing_postback_settings
WHERE url LIKE '%{offer_id}%'
)
AND (
url LIKE '%offer_id=&%'
OR url LIKE '%offer_id=$'
OR url REGEXP 'offer_id=[^a-zA-Z0-9]'
)
ORDER BY app_id, created_at;
Resend procedure:
- Run the detection query to get the list of affected
conversion_ids andtransaction_ids. - For each transaction, fetch the correct
offer_idfrom the PlayerAPIGET /api/v1/transactions/{txId}. - Feed the corrected data through
ResendWebhookService(app/Services/ResendWebhookService.php). The service already logs resends to theoutgoing-postbackAdActionEvent. - Verify resend success by checking status codes in Redshift for the new attempts.
Datadog monitor:
Create a scheduled query monitor (or log-based metric if AdActionEvent data flows to Datadog Logs) that fires when outgoing postbacks contain empty critical macros. Group by app_id. Thresholds:
- Warn: > 5 empty-macro postbacks in a 15-minute window for a single app.
- Critical: > 1% of total postbacks for an app have empty critical macros over a 1-hour window.
Layer C: Per-App Macro Canary Metric
Where: Same location as Layer A, inside V2Webhook::build() and V3Webhook::build().
What it emits:
For each macro that resolves to empty, increment a Datadog counter:
app(DatadogAdapter::class)->increment('outgoing_postback.empty_macro', [
'app_id' => $offerConverted->app_id,
'macro' => $macroName,
]);
This covers all macros, not just the critical set from Layer A.
Datadog anomaly monitor:
Set up an anomaly detection monitor on outgoing_postback.empty_macro grouped by app_id and macro. Use Datadog's built-in anomaly algorithm with:
- Algorithm: Agile (responds quickly to sudden changes).
- Deviations: 3 (avoids noise from minor fluctuations).
- Minimum data window: 24 hours (prevents false positives for new apps that don't have a baseline yet).
The monitor fires when a macro that historically resolves to a value suddenly starts showing up as empty. This catches regressions for any macro, not just the ones we've been burned by.
Relationship to Layer A:
Layers A and C both emit metrics from the same code location, but serve different purposes:
- Layer A's
empty_critical_macrometric has a static threshold monitor. Ifoffer_idis empty, alert immediately. - Layer C's
empty_macrometric has an anomaly monitor. If any macro's empty rate deviates from its historical baseline, alert.
Layer A catches known-important macros. Layer C catches unknown regressions.
Testing Strategy
- Unit tests: Verify that
V2Webhook::build()andV3Webhook::build()log warnings and emit metrics when critical macros resolve to empty. - Unit tests: Verify that the canary metric fires for any empty macro, not just critical ones.
- Integration tests: End-to-end test that mocks the PlayerAPI response with a null
offerId, flows throughOfferConverted::createFromRewardViaPlayerApi, and asserts the warning log and Datadog metric are emitted. - Redshift query validation: Run the detection query against staging data with known-empty macro postbacks to verify it returns the expected rows.
Rollout
- Ship Layer A (build-time validation) first. Low risk, immediate value.
- Set up the Layer B Datadog monitor and document the resend runbook.
- Ship Layer C (canary metric) and configure the anomaly monitor.
- Run the Layer B detection query against the 2026-03-30 to 2026-03-31 incident window and execute the resend runbook for affected publishers.