Offer Sync: Add a Per-App Metric
When to use this
Use this when an experiment needs a new per-app metric to reach Prism's ranking, for
example epc_d7_per_app next to today's rpc_d7_per_app. The dbt model already has the
per-app table; what this runbook covers is getting one more real per-app column through
the offer sync into the Offer API without re-creating the 2026-08-25 PATCH storm, and
without failing every metrics PATCH on the way.
Background on the two sync paths is in Offer Sync Today. The decision to rank per app is ADR 0071. This document is the procedure, not the decision.
On 2026-08-25 rpc_d7_per_app shipped on both sides 21 seconds apart, and the two sides
disagreed about when the key exists: the Offer API returned it on every offer, the
dashboard sent it only for some. The batch sync is a GET / diff / PATCH loop, so a key
present on one side only is a diff on every offer, every slot, forever. About 51,000
PATCH /v2/offers per 30 minutes (roughly 1,700 a minute) for 21 hours, almost none of
them writing anything (PUB-660). The only
stable shape is the key present on both sides, on every offer, null when there is no
value. Everything below exists to reach that shape without passing through a one-sided
state, in the one deploy order that works.
Prerequisites
- Write access to
AdGem/offer-api,AdGem/new_dashboardandAdAction/airflow-dags. - Edit access to the
adgem-dashboard-productionElastic Beanstalk environment properties (AWS console oraws elasticbeanstalk update-environment). That is whereOFFER_API_IGNORED_FIELDSlives: it is an EB environment property, not Parameter Store. The predeploy hook (.platform/hooks/predeploy/01-build-application-env-files.sh) merges EB properties into.envon every deploy. - Metabase access to the
Adgem Analysisdatabase for verification. - Datadog access.
What each hop knows about a per-app key
Verified against four repos during the review of this runbook on 2026-10-02: offer-api
main @ ae509c7 (released as 1.283.2 on 2026-10-05), new_dashboard master @
2552c8e (1.466.3), adgem/common 6.57.0, airflow-dags main @ c07e828,
targeted-api main @ 96691f7. Re-checked 2026-10-06. The column names below use
rpc_d7 as the example; a new metric follows the same route.
| Hop | Code | What it does with a key it does not know |
|---|---|---|
dbt metrics.app_campaign_metrics | airflow-dags/dbt/dbt-adgem/models/metrics/app_campaign_metrics.sql | Computes rpc_d7 per (app_id, campaign_id) by empirical-Bayes shrinkage toward a prior of COALESCE(NULLIF(campaign rpc_d7, 0), NULLIF(predicted_rpc, 0), 0). The other 18 metric columns (rpc_d1 through predicted_rpc) mirror campaign_metrics verbatim, plus rpc_d7_network: they exist so the row is shape-complete, they are not per-app numbers yet. |
Airflow sync_campaign_metrics | airflow-dags/lib/adgem/campaign_metrics.py (get_app_data_from_dbt, get_apps_by_campaign), adgem/ops/send_metrics_to_dash.py | The per-app query selects only campaign_id, app_id, rpc_d7; send_metrics_to_dash.py batches (100 campaigns or 2,000 app rows, whichever comes first) and posts to v1/campaigns/metrics. A new column has to be selected here or it never leaves the warehouse. |
| new_dashboard metrics endpoint | DTO app/Data/AppCampaignMetricData.php; upsert and prune in app/Data/CampaignMetricData.php (syncAppMetrics, reached from MetricController::store) | The DTO declares the 18 columns, every one except rpc_d7 as float $x = 0.0, and the app_campaign_metrics columns are NOT NULL DEFAULT 0. A column Airflow does not send is stored as 0.0, which Prism later reads as a real zero. Undeclared keys are ignored. Pruning runs only when the payload has apps: absent or null leaves the rows alone, [] prunes them all. |
| new_dashboard → Offer API, EventBridge path | app/Jobs/EventBridge/PublishCampaignMetricsChangedEventJob.php → offer-api app/EventSync/OfferSyncService.php:210-216 and app/EventSync/Handlers/CampaignMetricsChangedHandler.php | The event carries app_offer_metrics: {app_id: {rpc_d7_per_app}}; the handler folds the app's slice into offer_metrics, and OfferSyncService drops keys that are not in OfferMetric::$fillable, silently. No diff is involved: forward-compatible in both directions. The handler's absent-app default is ['rpc_d7_per_app' => null], naming that key only. |
| new_dashboard → Offer API, batch path | app/Libraries/OfferApi/Offers/OfferSyncFormatter.php → app/Helpers/OfferSyncHelper.php → PATCH /v2/offers/{id} | The formatter builds the whole offer, determinePatches diffs it against what the Offer API returned, and any difference inside offer_metrics sends the whole object. A key the Offer API returns and the dashboard omits is a permanent diff. A key the dashboard sends and the Offer API does not know is a 422 on the whole PATCH (OfferMetricsRules validates offer_metrics as array:<fillable>), real changes to the other metrics included. |
| offer-api | app/Models/OfferMetric.php, app/Libraries/Validation/Rules/OfferMetricsRules.php, app/Services/OfferService.php, app/OpenApi/PublicApi.php | $fillable decides what is stored and what validates, $visible what /v2/offers returns, $casts the JSON type, OfferMetricsRules the rule per key (numeric; rpc_d7_per_app is nullable|numeric), OfferService what stats exposes to Prism, and OfferMetric::carriesNetworkMetrics which keys cannot create a row on their own (it names rpc_d7_per_app only). |
| targeted-api (Prism) | app/Models/Enums/MetricsOfferSortFields.php, app/Libraries/OfferSorting/Sorters/MetricsOfferSorter.php | candidates() lists which stats keys a sort reads, in fallback order, per offer. Only matters if the new metric is to rank. |
Three consequences:
- EventBridge tolerates any deploy order. The dashboard can emit a key offer-api does not store yet (dropped), and offer-api can store a key the dashboard does not send yet (nothing arrives).
- The batch diff tolerates a one-sided key in neither direction, and the two failures
differ. Offer API first, dashboard omitting: the second loop of
determinePatchesnulls the key on every slot, the 08-25 loop. Dashboard first, Offer API unaware: every metrics PATCH for every app on the batch path is rejected with 422, so no metric syncs at all until offer-api catches up. The ignore list protects against the first, because an ignored key is never diffed on its own. It does not protect against the second, because the ignored key still rides insideoffer_metricswhenever another metric changes, and the validation rejects it on arrival. - So the order is forced: bridge the key on the ignore list, deploy offer-api, deploy the dashboard, then remove the bridge. The dashboard never ships ahead of offer-api.
Steps
1. Bridge the key on OFFER_API_IGNORED_FIELDS
Read the live value first; other keys live on that list and a pasted copy drops them:
aws elasticbeanstalk describe-configuration-settings \
--application-name adgem-dashboard --environment-name adgem-dashboard-production \
--profile adgem-legacy --region us-east-2 \
--query "ConfigurationSettings[0].OptionSettings[?OptionName=='OFFER_API_IGNORED_FIELDS'].Value" --output text
Append ,<key> to that value in the environment properties. Comma-separated, no
spaces. The EB aws:elasticbeanstalk:application:environment block is at its 4,096-byte
cap (the predeploy hook documents it, PUB-765), so a long key can fail to save; a short
one fits. The change restarts the app servers, Horizon workers included; a sync job cut
mid-run is picked up by the next slot. Afterwards, re-run the command above and check the
value took.
While a key is on that list, determinePatches skips it in both loops: it still travels
inside offer_metrics whenever another metric changes, but it can never produce a diff on
its own. From this moment the batch path cannot loop on that key, whatever either side
does. This is the lever that stopped the 08-25 storm, used as a planned bridge instead of
an emergency brake.
Two side effects of the bridge, both temporary:
- A change to that key alone never syncs via batch. It reaches the Offer API only when another metric on the same offer changes, or via EventBridge.
- The same list feeds
eventsync:diff --reconcile(app/Console/Commands/EventSyncDiff.php), so the drift report and the reconciler are blind to the key until step 5.
2. offer-api
One PR. Nothing here is gated on the dashboard, because the key is bridged.
- Migration:
offer_metrics.<key>, nullable, with the same precision as the dashboard'sapp_campaign_metricscolumn, which is defined inadgem/common(2026_08_13_120000_create_app_campaign_metrics_table.php):decimal(8, 4)forrpc_*andepc_*,decimal(12, 4)forrpm_*. Any difference means the dashboard sends1.31894567, offer-api stores1.3189, and the diff never converges. OfferMetric: add the key to$fillable,$visible,$casts(float, so it serialises as a JSON number) and the@propertydocblock, and add it to the keyscarriesNetworkMetrics()excludes. Without that, a payload carrying only{<key>: null}on an offer without a metrics row looks like it carries network metrics, andupdateOrCreatehits the NOT NULL network columns: a 500 instead of the guarded 200.OfferMetricsRules:nullable|numericfor the key. The defaultnumericrejects null, and null is how the dashboard says "no value for this app".OfferService::formatOfferForResponse: expose it instats, always present, null until the metric lands. Update the three PHPDoc response shapes (<key>: ?float), the OpenAPI property and thestatsobject'srequiredlist inPublicApi.php, and regenerate the committed file withcomposer openapi:generate:public. That last step was missed forrpc_d7_per_app(PEX-724).CampaignMetricsChangedHandler::patchForOffer: the absent-app default is['rpc_d7_per_app' => null]. Add the new key, or derive the list. Otherwise an app that drops out of the per-app table clearsrpc_d7_per_appbut keeps the new key's stale value, and realtime apps never pass through the batch sync to correct it.
Tests to add, with rpc_d7_per_app's as templates: a PATCH with the key as null
returns 200 and clears the value (tests/Feature/Offers/v2/UpdateOfferTest.php); the key
is present and null in /v1/offers stats when the row has no value
(tests/Feature/Offers/RetrieveActiveOffersByAppIdTest.php).
Deploy, and confirm the release is in production before step 4. Merged is not
enough: step 4 makes the dashboard send the key, and until offer-api's validation knows
it every metrics PATCH 422s. Note that /v1/offers is response-cached
(offer-cache.response_ttl, 1,200 seconds by default), so the new stats key shows per
app only once the cache turns over.
Precedent: offer-api #1467 (column, 2026-08-25), #1472 (stats and OpenAPI, 2026-08-31),
#1513 (accept null on PATCH, 2026-10-01), #1518 (always exposed, released 2026-10-05).
3. airflow-dags
Two things, and they are not optional:
- The dbt column has to be a real per-app number. Today only
rpc_d7is; the rest of the table mirrors the network value, so wiring a mirrored column through sends the network number under a per-app name. campaign_metrics.py: add the column to the per-app query and to the payloadget_apps_by_campaignbuilds, on every row. A row that arrives without the column is stored as0.0by the dashboard (next step), and0.0is a real zero to Prism, not a fallback.
The push is daily (@daily, after the dbt run) and can be triggered by hand; the
metrics-endpoint half is PEX-646 (one validation query per batch, bulk upsert).
4. new_dashboard
One PR. Only after step 2 is in production.
OfferSyncFormatter::formatOfferMetrics: always set the key: the value when the app has anapp_campaign_metricsrow, null when it does not. Never omit it, and never gate it on a flag. Omitting is what created the 08-25 asymmetry; a flag gate is what PEX-683 had to remove.AppCampaignMetricDataand theadgem/commonmigration: either guarantee step 3 sends a real value on every row, or make the column nullable with a?float $x = nullDTO property and let the formatter pass null through. Do not leave the0.0default in place for a column that may be absent.PublishCampaignMetricsChangedEventJob::appOfferMetrics: add the key to the per-app map next torpc_d7_per_app.OfferSyncHelper::nullDiffers: today it is pinned torpc_d7_per_app. Make it match the_per_appsuffix before adding a second metric, ornulland0will compare equal for the new one (PHP's loose==) and a null-to-zero change will never patch (PEX-673).- Fixtures: every faked Offer API offer in tests must carry the key (null), because that
is what offer-api returns from step 2 on.
UpdateOfferApiForAppTestandOfferSyncFormatterTestare the ones that broke last time.
Deploy. Precedent: new_dashboard #2872 (2026-08-25), #2920 (2026-10-01), #2925 (1.466.3, 2026-10-02).
5. Wait one nightly refresh, then remove the bridge
Wait for the nightly metrics refresh after step 4 (the push lands around 00:10 UTC and the batch slots carry it until about 01:10). During it every offer whose other metrics moved takes the new key along, so by morning almost every offer with a per-app row already holds the value. Removing the entry before that refresh makes the first slot patch every offer with a per-app row at once, and burst rate is what trips the 2,000-a-minute limit.
Then remove the key from OFFER_API_IGNORED_FIELDS the same way it was added: read the
live value, remove ,<key>, save, let the environment restart, confirm the value took.
The first slot after that patches only offers whose stored value still differs from the
dashboard's, once; after that only real changes patch, and the drift report sees the key
again.
6. Watch for a day
sum:trace.laravel.request.hits{service:offer-api,env:production,resource_name:*v2.offers.update*}.as_count()in 5-minute buckets. Baseline in business hours is 10 to 80 per 5 minutes. The nightly refresh runs from about 00:10 to 01:10 UTC and peaks around 7,000 per 5 minutes (44,960 in the 00:00 hour and 11,636 in the 01:00 hour on 2026-10-02), with a smaller bump of 2,500 to 3,400 in the 03:00 hour. That peak is roughly the storm's rate, so duration is what tells them apart: the refresh is back to baseline within the hour; the 08-25 storm held about 1,700 a minute, flat, for 21 hours.offer_api.eventsync.consume.offer_metrics_field_diff{field:<key>}, taggedkind:valueorkind:precision: how often an incoming EventSync value differed from the stored one before the write (the PUB-830 pre-flight diff). It counts only keys already in$fillable, only offers with a metrics row loaded, and emits nothing for null over null, so a zero does not prove the realtime cohort receives the key; a non-zero proves it does.- The requests-to-writes ratio, if in doubt. PUB-660's timeline put the storm at roughly 300,000 requests for about 75 metric-row writes over the sampled window. A healthy sync writes roughly 1:1.
Verification
The Offer API data in Redshift arrives as a materialized view,
adgem_raw.offer_api_mirror.offer_metrics (Liquibase changeset 28), which freezes its
column list at creation; REFRESH refreshes rows, not schema. The offer_api.offer_metrics
table that Metabase queries is a dbt SELECT * over that view
(dbt/dbt-offer-api/models/offer_api/offer_metrics.sql), so it inherits the frozen list. A
new column is invisible in Metabase until someone recreates the view with a changeset in
airflow-dags/migrations/changelog-adgem_raw.yaml; changeset 46 did this for
rpc_d7_per_app (PEX-654, airflow-dags #2689). Do that early, or you cannot verify from
Metabase at all.
Once the column exists, per pilot app:
SELECT o.adgem_app_id AS app_id,
SUM(CASE WHEN m.campaign_id IS NOT NULL THEN 1 ELSE 0 END) AS dbt_rows,
SUM(CASE WHEN m.campaign_id IS NOT NULL
AND COALESCE(ROUND(m.<dbt_column>, 4) = om.<key>, FALSE) THEN 1 ELSE 0 END) AS in_sync,
SUM(CASE WHEN m.campaign_id IS NOT NULL AND om.<key> IS NULL THEN 1 ELSE 0 END) AS missing_in_offer_api,
SUM(CASE WHEN m.campaign_id IS NULL AND om.<key> IS NULL THEN 1 ELSE 0 END) AS null_on_both_sides
FROM offer_api.offers o
LEFT JOIN offer_api.offer_metrics om ON om.offer_id = o.id
LEFT JOIN metrics.app_campaign_metrics m ON m.app_id = o.adgem_app_id AND m.campaign_id = o.campaign_id
WHERE o.adgem_app_id IN (<pilot apps>) AND o.disabled_at IS NULL AND o.deleted_at IS NULL
GROUP BY 1;
Reading it:
missing_in_offer_apishould reach zero within one slot of the app's batch bucket after step 5 (twiceat:10and:40,once20at:20,once50at:50, with apps staggered 20 seconds apart, so the last app of a slot can run about 15 minutes after it starts), or within minutes on the realtime cohort. The mirror lags up to 30 minutes plus theoffer_api_to_aarun time.- An offer with no metrics row at all returns
offer_metrics: []from/v2/offers, so "present on every offer" holds for offers with a row; those offers show asnull_on_both_sideshere, which is correct. - dbt rebuilds
app_campaign_metricsshortly after 00:00 UTC, before the push lands, so running this between the rebuild and about 01:10 shows false mismatches. Run it later in the morning.
At serve time, the only direct signal is Prism's
targeted_api.offer_sorting.per_app_rpc{adgem_app_id, fell_back} (today specific to
rpc_d7_per_app): fell_back:false means the per-app value was present for that offer.
Traps
- A new key has no flag to lean on; the ignore entry is its only bridge. PEX-683 removed the per-app data flags precisely so that no flag state can recreate the asymmetry. Do not skip step 1.
- Never let the dashboard ship the key before offer-api is in production. That is not a slow loop, it is every metrics PATCH on the batch path failing with 422.
- Do not drop the key from
$visibleas a "cheap fix" if the sync ever diffs on it. That converges only while the dashboard never sends the key, and re-arms the loop the moment it does. nulland0mean different things to Prism. Null falls back tonetwork_epcfor that offer (MetricsOfferSortFields::candidates()); 0 is a real zero, which in the default descending sort ranks the offer last, tied with any fallback offer whosenetwork_epcis 0. Keep that innullDiffers, in the formatter (null without a row, never0.0), in the DTO defaults, and in any analysis.- The dashboard's per-app row set is a snapshot. Apps missing from an Airflow push are pruned, and the formatter then sends null, which clears the value in offer-api. That is correct, not a bug: a cell that dropped out of the per-app table has no per-app value. The realtime path needs the handler default from step 2.5 to do the same.
- The debug log that would have named the 08-25 bug is discarded in production.
OfferSyncHelperlogs the parent key (offer_metrics), the computed diff and both sides atLog::debug; production runs atSINGLE_LOG_LEVEL=warning. If a sync misbehaves, read the request-to-write ratio and the per-field EventSync metric instead of waiting for a log line.
Status and open items, as of 2026-10-06
- PEX-571 (the key on every offer for flagged
apps) and PEX-683 (the key on every offer,
no flags) are complete: new_dashboard 1.466.3 (2026-10-02 16:56 UTC), offer-api 1.283.2
(2026-10-05), the
rpc_d7_per_appentry removed fromOFFER_API_IGNORED_FIELDS, anduses-per-app-rpcanduses-per-app-rpc-ingestiondeleted in DevCycle (reported done 2026-10-06).IgnoredFieldsis gone; the configured list goes straight todeterminePatches. - PEX-673:
nullDiffersexists but is pinned to one key (step 4.4). - PEX-724:
openapi.public.yamlstill lacksrpc_d7_per_app; two docblocks (offer-apicarriesNetworkMetrics, targeted-apicandidates()) describe the pre-#1518 behaviour. - There is still no round-trip contract test: format an offer with
OfferSyncFormatter, pass it through the shapeOfferResourcereturns (including theoffer_metrics: []case for offers without a row), assertdeterminePatchesis empty. Nothing in either repo fails today when the two shapes drift apart, and PR review of either half alone cannot catch it. It is the single most useful thing to add before the next metric.