Authenticating...
Skip to main content

Offer Sync: Add a Per-App Metric

When to use this​

Use this when an experiment needs a new per-app metric to reach Prism's ranking, for example epc_d7_per_app next to today's rpc_d7_per_app. The dbt model already has the per-app table; what this runbook covers is getting one more real per-app column through the offer sync into the Offer API without re-creating the 2026-08-25 PATCH storm, and without failing every metrics PATCH on the way.

Background on the two sync paths is in Offer Sync Today. The decision to rank per app is ADR 0071. This document is the procedure, not the decision.

The storm was a shape problem, not a volume problem

On 2026-08-25 rpc_d7_per_app shipped on both sides 21 seconds apart, and the two sides disagreed about when the key exists: the Offer API returned it on every offer, the dashboard sent it only for some. The batch sync is a GET / diff / PATCH loop, so a key present on one side only is a diff on every offer, every slot, forever. About 51,000 PATCH /v2/offers per 30 minutes (roughly 1,700 a minute) for 21 hours, almost none of them writing anything (PUB-660). The only stable shape is the key present on both sides, on every offer, null when there is no value. Everything below exists to reach that shape without passing through a one-sided state, in the one deploy order that works.

Prerequisites​

  • Write access to AdGem/offer-api, AdGem/new_dashboard and AdAction/airflow-dags.
  • Edit access to the adgem-dashboard-production Elastic Beanstalk environment properties (AWS console or aws elasticbeanstalk update-environment). That is where OFFER_API_IGNORED_FIELDS lives: it is an EB environment property, not Parameter Store. The predeploy hook (.platform/hooks/predeploy/01-build-application-env-files.sh) merges EB properties into .env on every deploy.
  • Metabase access to the Adgem Analysis database for verification.
  • Datadog access.

What each hop knows about a per-app key​

Verified against four repos during the review of this runbook on 2026-10-02: offer-api main @ ae509c7 (released as 1.283.2 on 2026-10-05), new_dashboard master @ 2552c8e (1.466.3), adgem/common 6.57.0, airflow-dags main @ c07e828, targeted-api main @ 96691f7. Re-checked 2026-10-06. The column names below use rpc_d7 as the example; a new metric follows the same route.

HopCodeWhat it does with a key it does not know
dbt metrics.app_campaign_metricsairflow-dags/dbt/dbt-adgem/models/metrics/app_campaign_metrics.sqlComputes rpc_d7 per (app_id, campaign_id) by empirical-Bayes shrinkage toward a prior of COALESCE(NULLIF(campaign rpc_d7, 0), NULLIF(predicted_rpc, 0), 0). The other 18 metric columns (rpc_d1 through predicted_rpc) mirror campaign_metrics verbatim, plus rpc_d7_network: they exist so the row is shape-complete, they are not per-app numbers yet.
Airflow sync_campaign_metricsairflow-dags/lib/adgem/campaign_metrics.py (get_app_data_from_dbt, get_apps_by_campaign), adgem/ops/send_metrics_to_dash.pyThe per-app query selects only campaign_id, app_id, rpc_d7; send_metrics_to_dash.py batches (100 campaigns or 2,000 app rows, whichever comes first) and posts to v1/campaigns/metrics. A new column has to be selected here or it never leaves the warehouse.
new_dashboard metrics endpointDTO app/Data/AppCampaignMetricData.php; upsert and prune in app/Data/CampaignMetricData.php (syncAppMetrics, reached from MetricController::store)The DTO declares the 18 columns, every one except rpc_d7 as float $x = 0.0, and the app_campaign_metrics columns are NOT NULL DEFAULT 0. A column Airflow does not send is stored as 0.0, which Prism later reads as a real zero. Undeclared keys are ignored. Pruning runs only when the payload has apps: absent or null leaves the rows alone, [] prunes them all.
new_dashboard → Offer API, EventBridge pathapp/Jobs/EventBridge/PublishCampaignMetricsChangedEventJob.php → offer-api app/EventSync/OfferSyncService.php:210-216 and app/EventSync/Handlers/CampaignMetricsChangedHandler.phpThe event carries app_offer_metrics: {app_id: {rpc_d7_per_app}}; the handler folds the app's slice into offer_metrics, and OfferSyncService drops keys that are not in OfferMetric::$fillable, silently. No diff is involved: forward-compatible in both directions. The handler's absent-app default is ['rpc_d7_per_app' => null], naming that key only.
new_dashboard → Offer API, batch pathapp/Libraries/OfferApi/Offers/OfferSyncFormatter.php → app/Helpers/OfferSyncHelper.php → PATCH /v2/offers/{id}The formatter builds the whole offer, determinePatches diffs it against what the Offer API returned, and any difference inside offer_metrics sends the whole object. A key the Offer API returns and the dashboard omits is a permanent diff. A key the dashboard sends and the Offer API does not know is a 422 on the whole PATCH (OfferMetricsRules validates offer_metrics as array:<fillable>), real changes to the other metrics included.
offer-apiapp/Models/OfferMetric.php, app/Libraries/Validation/Rules/OfferMetricsRules.php, app/Services/OfferService.php, app/OpenApi/PublicApi.php$fillable decides what is stored and what validates, $visible what /v2/offers returns, $casts the JSON type, OfferMetricsRules the rule per key (numeric; rpc_d7_per_app is nullable|numeric), OfferService what stats exposes to Prism, and OfferMetric::carriesNetworkMetrics which keys cannot create a row on their own (it names rpc_d7_per_app only).
targeted-api (Prism)app/Models/Enums/MetricsOfferSortFields.php, app/Libraries/OfferSorting/Sorters/MetricsOfferSorter.phpcandidates() lists which stats keys a sort reads, in fallback order, per offer. Only matters if the new metric is to rank.

Three consequences:

  • EventBridge tolerates any deploy order. The dashboard can emit a key offer-api does not store yet (dropped), and offer-api can store a key the dashboard does not send yet (nothing arrives).
  • The batch diff tolerates a one-sided key in neither direction, and the two failures differ. Offer API first, dashboard omitting: the second loop of determinePatches nulls the key on every slot, the 08-25 loop. Dashboard first, Offer API unaware: every metrics PATCH for every app on the batch path is rejected with 422, so no metric syncs at all until offer-api catches up. The ignore list protects against the first, because an ignored key is never diffed on its own. It does not protect against the second, because the ignored key still rides inside offer_metrics whenever another metric changes, and the validation rejects it on arrival.
  • So the order is forced: bridge the key on the ignore list, deploy offer-api, deploy the dashboard, then remove the bridge. The dashboard never ships ahead of offer-api.

Steps​

1. Bridge the key on OFFER_API_IGNORED_FIELDS​

Read the live value first; other keys live on that list and a pasted copy drops them:

aws elasticbeanstalk describe-configuration-settings \
--application-name adgem-dashboard --environment-name adgem-dashboard-production \
--profile adgem-legacy --region us-east-2 \
--query "ConfigurationSettings[0].OptionSettings[?OptionName=='OFFER_API_IGNORED_FIELDS'].Value" --output text

Append ,<key> to that value in the environment properties. Comma-separated, no spaces. The EB aws:elasticbeanstalk:application:environment block is at its 4,096-byte cap (the predeploy hook documents it, PUB-765), so a long key can fail to save; a short one fits. The change restarts the app servers, Horizon workers included; a sync job cut mid-run is picked up by the next slot. Afterwards, re-run the command above and check the value took.

While a key is on that list, determinePatches skips it in both loops: it still travels inside offer_metrics whenever another metric changes, but it can never produce a diff on its own. From this moment the batch path cannot loop on that key, whatever either side does. This is the lever that stopped the 08-25 storm, used as a planned bridge instead of an emergency brake.

Two side effects of the bridge, both temporary:

  • A change to that key alone never syncs via batch. It reaches the Offer API only when another metric on the same offer changes, or via EventBridge.
  • The same list feeds eventsync:diff --reconcile (app/Console/Commands/EventSyncDiff.php), so the drift report and the reconciler are blind to the key until step 5.

2. offer-api​

One PR. Nothing here is gated on the dashboard, because the key is bridged.

  1. Migration: offer_metrics.<key>, nullable, with the same precision as the dashboard's app_campaign_metrics column, which is defined in adgem/common (2026_08_13_120000_create_app_campaign_metrics_table.php): decimal(8, 4) for rpc_* and epc_*, decimal(12, 4) for rpm_*. Any difference means the dashboard sends 1.31894567, offer-api stores 1.3189, and the diff never converges.
  2. OfferMetric: add the key to $fillable, $visible, $casts (float, so it serialises as a JSON number) and the @property docblock, and add it to the keys carriesNetworkMetrics() excludes. Without that, a payload carrying only {<key>: null} on an offer without a metrics row looks like it carries network metrics, and updateOrCreate hits the NOT NULL network columns: a 500 instead of the guarded 200.
  3. OfferMetricsRules: nullable|numeric for the key. The default numeric rejects null, and null is how the dashboard says "no value for this app".
  4. OfferService::formatOfferForResponse: expose it in stats, always present, null until the metric lands. Update the three PHPDoc response shapes (<key>: ?float), the OpenAPI property and the stats object's required list in PublicApi.php, and regenerate the committed file with composer openapi:generate:public. That last step was missed for rpc_d7_per_app (PEX-724).
  5. CampaignMetricsChangedHandler::patchForOffer: the absent-app default is ['rpc_d7_per_app' => null]. Add the new key, or derive the list. Otherwise an app that drops out of the per-app table clears rpc_d7_per_app but keeps the new key's stale value, and realtime apps never pass through the batch sync to correct it.

Tests to add, with rpc_d7_per_app's as templates: a PATCH with the key as null returns 200 and clears the value (tests/Feature/Offers/v2/UpdateOfferTest.php); the key is present and null in /v1/offers stats when the row has no value (tests/Feature/Offers/RetrieveActiveOffersByAppIdTest.php).

Deploy, and confirm the release is in production before step 4. Merged is not enough: step 4 makes the dashboard send the key, and until offer-api's validation knows it every metrics PATCH 422s. Note that /v1/offers is response-cached (offer-cache.response_ttl, 1,200 seconds by default), so the new stats key shows per app only once the cache turns over.

Precedent: offer-api #1467 (column, 2026-08-25), #1472 (stats and OpenAPI, 2026-08-31), #1513 (accept null on PATCH, 2026-10-01), #1518 (always exposed, released 2026-10-05).

3. airflow-dags​

Two things, and they are not optional:

  1. The dbt column has to be a real per-app number. Today only rpc_d7 is; the rest of the table mirrors the network value, so wiring a mirrored column through sends the network number under a per-app name.
  2. campaign_metrics.py: add the column to the per-app query and to the payload get_apps_by_campaign builds, on every row. A row that arrives without the column is stored as 0.0 by the dashboard (next step), and 0.0 is a real zero to Prism, not a fallback.

The push is daily (@daily, after the dbt run) and can be triggered by hand; the metrics-endpoint half is PEX-646 (one validation query per batch, bulk upsert).

4. new_dashboard​

One PR. Only after step 2 is in production.

  1. OfferSyncFormatter::formatOfferMetrics: always set the key: the value when the app has an app_campaign_metrics row, null when it does not. Never omit it, and never gate it on a flag. Omitting is what created the 08-25 asymmetry; a flag gate is what PEX-683 had to remove.
  2. AppCampaignMetricData and the adgem/common migration: either guarantee step 3 sends a real value on every row, or make the column nullable with a ?float $x = null DTO property and let the formatter pass null through. Do not leave the 0.0 default in place for a column that may be absent.
  3. PublishCampaignMetricsChangedEventJob::appOfferMetrics: add the key to the per-app map next to rpc_d7_per_app.
  4. OfferSyncHelper::nullDiffers: today it is pinned to rpc_d7_per_app. Make it match the _per_app suffix before adding a second metric, or null and 0 will compare equal for the new one (PHP's loose ==) and a null-to-zero change will never patch (PEX-673).
  5. Fixtures: every faked Offer API offer in tests must carry the key (null), because that is what offer-api returns from step 2 on. UpdateOfferApiForAppTest and OfferSyncFormatterTest are the ones that broke last time.

Deploy. Precedent: new_dashboard #2872 (2026-08-25), #2920 (2026-10-01), #2925 (1.466.3, 2026-10-02).

5. Wait one nightly refresh, then remove the bridge​

Wait for the nightly metrics refresh after step 4 (the push lands around 00:10 UTC and the batch slots carry it until about 01:10). During it every offer whose other metrics moved takes the new key along, so by morning almost every offer with a per-app row already holds the value. Removing the entry before that refresh makes the first slot patch every offer with a per-app row at once, and burst rate is what trips the 2,000-a-minute limit.

Then remove the key from OFFER_API_IGNORED_FIELDS the same way it was added: read the live value, remove ,<key>, save, let the environment restart, confirm the value took. The first slot after that patches only offers whose stored value still differs from the dashboard's, once; after that only real changes patch, and the drift report sees the key again.

6. Watch for a day​

  • sum:trace.laravel.request.hits{service:offer-api,env:production,resource_name:*v2.offers.update*}.as_count() in 5-minute buckets. Baseline in business hours is 10 to 80 per 5 minutes. The nightly refresh runs from about 00:10 to 01:10 UTC and peaks around 7,000 per 5 minutes (44,960 in the 00:00 hour and 11,636 in the 01:00 hour on 2026-10-02), with a smaller bump of 2,500 to 3,400 in the 03:00 hour. That peak is roughly the storm's rate, so duration is what tells them apart: the refresh is back to baseline within the hour; the 08-25 storm held about 1,700 a minute, flat, for 21 hours.
  • offer_api.eventsync.consume.offer_metrics_field_diff{field:<key>}, tagged kind:value or kind:precision: how often an incoming EventSync value differed from the stored one before the write (the PUB-830 pre-flight diff). It counts only keys already in $fillable, only offers with a metrics row loaded, and emits nothing for null over null, so a zero does not prove the realtime cohort receives the key; a non-zero proves it does.
  • The requests-to-writes ratio, if in doubt. PUB-660's timeline put the storm at roughly 300,000 requests for about 75 metric-row writes over the sampled window. A healthy sync writes roughly 1:1.

Verification​

The Offer API data in Redshift arrives as a materialized view, adgem_raw.offer_api_mirror.offer_metrics (Liquibase changeset 28), which freezes its column list at creation; REFRESH refreshes rows, not schema. The offer_api.offer_metrics table that Metabase queries is a dbt SELECT * over that view (dbt/dbt-offer-api/models/offer_api/offer_metrics.sql), so it inherits the frozen list. A new column is invisible in Metabase until someone recreates the view with a changeset in airflow-dags/migrations/changelog-adgem_raw.yaml; changeset 46 did this for rpc_d7_per_app (PEX-654, airflow-dags #2689). Do that early, or you cannot verify from Metabase at all.

Once the column exists, per pilot app:

SELECT o.adgem_app_id AS app_id,
SUM(CASE WHEN m.campaign_id IS NOT NULL THEN 1 ELSE 0 END) AS dbt_rows,
SUM(CASE WHEN m.campaign_id IS NOT NULL
AND COALESCE(ROUND(m.<dbt_column>, 4) = om.<key>, FALSE) THEN 1 ELSE 0 END) AS in_sync,
SUM(CASE WHEN m.campaign_id IS NOT NULL AND om.<key> IS NULL THEN 1 ELSE 0 END) AS missing_in_offer_api,
SUM(CASE WHEN m.campaign_id IS NULL AND om.<key> IS NULL THEN 1 ELSE 0 END) AS null_on_both_sides
FROM offer_api.offers o
LEFT JOIN offer_api.offer_metrics om ON om.offer_id = o.id
LEFT JOIN metrics.app_campaign_metrics m ON m.app_id = o.adgem_app_id AND m.campaign_id = o.campaign_id
WHERE o.adgem_app_id IN (<pilot apps>) AND o.disabled_at IS NULL AND o.deleted_at IS NULL
GROUP BY 1;

Reading it:

  • missing_in_offer_api should reach zero within one slot of the app's batch bucket after step 5 (twice at :10 and :40, once20 at :20, once50 at :50, with apps staggered 20 seconds apart, so the last app of a slot can run about 15 minutes after it starts), or within minutes on the realtime cohort. The mirror lags up to 30 minutes plus the offer_api_to_aa run time.
  • An offer with no metrics row at all returns offer_metrics: [] from /v2/offers, so "present on every offer" holds for offers with a row; those offers show as null_on_both_sides here, which is correct.
  • dbt rebuilds app_campaign_metrics shortly after 00:00 UTC, before the push lands, so running this between the rebuild and about 01:10 shows false mismatches. Run it later in the morning.

At serve time, the only direct signal is Prism's targeted_api.offer_sorting.per_app_rpc{adgem_app_id, fell_back} (today specific to rpc_d7_per_app): fell_back:false means the per-app value was present for that offer.

Traps​

  • A new key has no flag to lean on; the ignore entry is its only bridge. PEX-683 removed the per-app data flags precisely so that no flag state can recreate the asymmetry. Do not skip step 1.
  • Never let the dashboard ship the key before offer-api is in production. That is not a slow loop, it is every metrics PATCH on the batch path failing with 422.
  • Do not drop the key from $visible as a "cheap fix" if the sync ever diffs on it. That converges only while the dashboard never sends the key, and re-arms the loop the moment it does.
  • null and 0 mean different things to Prism. Null falls back to network_epc for that offer (MetricsOfferSortFields::candidates()); 0 is a real zero, which in the default descending sort ranks the offer last, tied with any fallback offer whose network_epc is 0. Keep that in nullDiffers, in the formatter (null without a row, never 0.0), in the DTO defaults, and in any analysis.
  • The dashboard's per-app row set is a snapshot. Apps missing from an Airflow push are pruned, and the formatter then sends null, which clears the value in offer-api. That is correct, not a bug: a cell that dropped out of the per-app table has no per-app value. The realtime path needs the handler default from step 2.5 to do the same.
  • The debug log that would have named the 08-25 bug is discarded in production. OfferSyncHelper logs the parent key (offer_metrics), the computed diff and both sides at Log::debug; production runs at SINGLE_LOG_LEVEL=warning. If a sync misbehaves, read the request-to-write ratio and the per-field EventSync metric instead of waiting for a log line.

Status and open items, as of 2026-10-06​

  • PEX-571 (the key on every offer for flagged apps) and PEX-683 (the key on every offer, no flags) are complete: new_dashboard 1.466.3 (2026-10-02 16:56 UTC), offer-api 1.283.2 (2026-10-05), the rpc_d7_per_app entry removed from OFFER_API_IGNORED_FIELDS, and uses-per-app-rpc and uses-per-app-rpc-ingestion deleted in DevCycle (reported done 2026-10-06). IgnoredFields is gone; the configured list goes straight to determinePatches.
  • PEX-673: nullDiffers exists but is pinned to one key (step 4.4).
  • PEX-724: openapi.public.yaml still lacks rpc_d7_per_app; two docblocks (offer-api carriesNetworkMetrics, targeted-api candidates()) describe the pre-#1518 behaviour.
  • There is still no round-trip contract test: format an offer with OfferSyncFormatter, pass it through the shape OfferResource returns (including the offer_metrics: [] case for offers without a row), assert determinePatches is empty. Nothing in either repo fails today when the two shapes drift apart, and PR review of either half alone cannot catch it. It is the single most useful thing to add before the next metric.