Offer Sync Today
How AdGem offers get from the dashboard into the Offer API, which the Prism API and offerwalls read from. This page covers the moving parts, who is on which sync path, what limits us, what we monitor, and the traps that have caused incidents. The last section covers where this is going.
Audience: anyone who touches offers, campaigns, app enrollment, or the Offer API, or who is on call for them.
Runbooks: Enroll an App and EventSync Incident Response.
State as of 2026-09-29. Counts and flag values change weekly. Each section says where to read the live value.
TL;DR
- The dashboard (
new_dashboard) is the source of truth for offers. The Offer API (offer-api) holds a copy per app, one row per app and campaign. - There are two main sync paths, and each app should be on exactly one of them. Nothing enforces that (see The flags):
- Batch sync: a scheduled job that recomputes an app's offers and pushes the difference over HTTP. 282 apps.
- Realtime sync (EventSync): dashboard changes publish events to EventBridge, and a consumer inside
offer-apiwrites them straight to its database. 103 apps.
- A differ runs every 30 minutes over the realtime apps. When it finds drift, the reconciler heals those apps through the batch job, at most 25 apps per run.
- Which path an app is on, and whether it gets Prism, is decided by three hand-edited DevCycle lists that nothing cross-checks. That gap caused IR-308.
- Realtime is close to capacity: at 103 apps, the differ uses 80-88% of its 25-minute lock.
The big picture
Two ideas explain most of the behaviour:
- Offers are mirrored per app. A campaign that runs on 190 apps is 190 offer rows. One campaign edit can cost up to about 190 offer writes. Bulk campaign work is what saturates the integration.
- Some fields are part of an offer's identity. Changing a contract field (
total_amount,total_payout_usd,goals,is_multi_reward,location_targets,device_targets,os_targets) supersedes the offer: the old row is soft-deleted and a new one created. Changing any other field is an in-place patch. The list lives inconfig/offer-api.php(contract_fields) innew_dashboard, andoffer-apikeeps its own copy for EventSync.
Path 1: batch sync
The original path. It is being retired (see Where this is going).
- Command:
offer-api:sync-v2 --frequency=<bucket>innew_dashboard, scheduled inapp/Console/Kernel.php. All runs useonOneServer(), and none usewithoutOverlapping():twiceat :10 and :40once20at :20once50at :50
- Who: the DevCycle JSON variable
offer-api-v2-sync-schedule, shaped{"once20": [...], "once50": [...], "twice": [...]}. It has 135 / 135 / 12 = 282 apps. Only apps withapp_status = activeare dispatched. - What it does: it dispatches
UpdateOfferApiForAppJob($appId)for each app, staggered 20 seconds apart, so one slot's apps spread across many minutes. The job then:- re-checks that the app is still active;
- builds the expected offer set from
AppCampaign(allJoins(),basicCampaignFilters(),OfferSyncFormatter::formatOffer), dropping offers withtotal_payout_usd == 0; - reads the app's current offers with
GET /v2/offers?app_id=(paginated); - supersedes (DELETE then POST) any offer whose contract fields changed, PATCHes other changes, and PATCHes
{"disabled": true}onto offers that should no longer be live.
- Queue:
default(Horizonsupervisor-1, 5 processes,tries: 3, Horizon's default 60-second timeout). The same queue carries EventSync publishes, transactional email, and Cognito registration. - Break-glass:
offer-api:sync-v2 --apps=1,2,3syncs specific apps now, ignoring the schedule. It still skips apps that aren't active.
The batch dispatcher does not exclude realtime apps. The two paths stay separate only because people move an app from one list to the other by hand.
Path 2: realtime sync (EventSync)
The target path for every app. ADR 0057 records the decision.
Publish (dashboard side). Model changes dispatch one of 11 publish jobs in app/Jobs/EventBridge/. Each job gates on the app being in the realtime list, and each sends one event type to the EventBridge bus (source adgem.dashboard):
| Event | Event | Event |
|---|---|---|
AppCampaignMembershipChanged | CampaignMetricsChanged | OfferwallFieldsChanged |
AppStatusChanged | CampaignStatusChanged | OfferwallMultiplierChanged |
CampaignCreativeChanged | GoalChanged | TargetingProfileChanged |
CampaignFieldsChanged | MarginOverrideChanged |
Publish jobs retry with backoff (10s, 30s, 60s). Events larger than 240 KB are refused (EVENTBRIDGE_MAX_DETAIL_BYTES).
Consume (Offer API side). EventBridge rules route events into two SQS queues, supersession and update. eventsync:consume --queue=<name> runs as a systemd service (eventsync-consumer@<queue>), one process per queue, restarting every 9 minutes (--max-time=540). Handlers in app/EventSync/Handlers/ apply each event through OfferSyncService, which writes directly to the Offer API database. There is no HTTP call, so realtime writes never touch the HTTP rate limit.
Who: the DevCycle JSON variable real-time-offer-sync-apps, shaped {"app_ids": [...]}, with 103 apps. The dashboard reads it through DevCycleService::realTimeSyncApps() and caches it for 30 seconds. If DevCycle fails or serves the default, the code keeps the last good list rather than emptying the cohort. To empty the cohort on purpose, set app_ids to []. Turning the feature off does not empty it.
No backfill. Events only carry changes made after an app joins. When an app joins, its first full sync comes from the reconciler (next section).
The differ and the reconciler
eventsync:diff in new_dashboard is the realtime path's safety net. It compares what the dashboard says an app should have against what the Offer API has, and can heal the difference.
- Schedule:
eventsync:diff --reconcile --shard=i/Nat :05 and :35,onOneServer(),withoutOverlapping(25).NisEVENTSYNC_DIFF_SHARDS(default 2). - Cohort: the realtime list (the same list the publish jobs use, so an app on the event path is always covered by the differ). It emits
cohort_size. - What each offer gets:
| Action | Meaning | Why it matters |
|---|---|---|
would_create | Expected, but absent from the Offer API | Missing inventory: the publisher can't show an offer it should have |
would_disable | Live in the Offer API, but not expected | A stale offer is being served. The most serious case |
would_supersede | On both sides, a contract field differs | Soft-delete plus recreate |
would_patch | On both sides, only other fields differ | In-place patch |
An extra offer that is already disabled is noop_already_disabled and isn't counted.
-
Reconcile: for an app with drift, it dispatches
UpdateOfferApiForAppJob(the batch job, over HTTP, on purpose: the heal path should not share the event path's failure modes). Reconcile is:- gated on the DevCycle flag
eventsync-reconcile-enabled; - limited to active apps;
- capped at 25 apps per run across all shards (
OFFER_API_RECONCILE_MAX_APPS_PER_RUN). Apps over the cap are skipped withreason:cap_exceededand wait for a later run. A claim rotation (DiffShardCoordinator) makes sure the same apps don't always win the slots.
- gated on the DevCycle flag
-
Manual run (read-only unless you add
--reconcile):php artisan eventsync:diff 29011 30035 --show-noopApp ids are positional. Exit code is 0 even when drift is found, and 1 only when an app errors.
The HTTP side paths
Two smaller paths remain:
- App enable/disable. When an app's
app_statuschanges,AppObserverdispatchesDisableOfferForAppJoborEnableOfferForAppJobon theapp-campaign-syncqueue, and always publishesAppStatusChanged. The jobs skip realtime apps (the event covers them). For batch apps this is the only thing that disables a whole inactive app, because the batch only visits active apps. - Location targets.
offer-api:sync-location-targetspushes countries, states and cities through the v1 API. It is run by hand; it isn't scheduled.
The flags
Flag (DevCycle, adgem project) | Type | Read by | What it decides |
|---|---|---|---|
real-time-offer-sync-apps | JSON {app_ids} | new_dashboard | Realtime cohort, and the differ's cohort |
offer-api-v2-sync-schedule | JSON buckets | new_dashboard | Batch cohort and cadence |
uses-targeted-api | Boolean, per app | api (offerwall) | Whether the offerwall is served from Prism |
eventsync-reconcile-enabled | Boolean | new_dashboard | Whether the differ heals or only reports |
The first three lists are independent, and nothing checks them against each other.
- An app in
uses-targeted-apibut on neither sync list has no offers in the Offer API, so Prism serves an empty wall. That was IR-308: 39 of 77 Prism-enabled apps had zero inventory. - An app in both sync lists is synced twice. An app removed from one list and not added to the other stops syncing silently.
- The batch schedule has other editors. Read it live before editing, and never paste a saved copy back: that deletes whatever was added since.
Is app X synced?
-
Which list is it in? Look for the app id in
real-time-offer-sync-apps(realtime) andoffer-api-v2-sync-schedule(batch).- Exactly one: it is on that sync path.
- Neither: nothing syncs it. It has no offers in the Offer API, or only stale ones.
- Both: it is synced twice. Remove it from one.
-
Is it eligible? Both paths skip apps that aren't
active, even if they are listed. An app can beactiveand still be soft-deleted, so checkdeleted_attoo. -
Does the Offer API match? On a dashboard instance, run the read-only differ for the app:
php artisan eventsync:diff <app_id>Its
Summary:line should show0 missingand0 extra actionable. This works for any app id, including batch apps. -
Prism? Check whether
uses-targeted-apiis on for the app. That is only safe once step 3 passes.
Limits today
| Component | Limit | Where it stands |
|---|---|---|
| Differ | Full diff of every realtime app every 30 minutes, inside a 25-minute lock | 80-88% of the lock at 103 apps (measured 2026-09-23). The biggest blocker to growing realtime |
| Consumer | One process per queue. Cost grows linearly with cohort size, about 3.25 ms per app per message | Drains a nightly burst in about 10 minutes at today's size |
| Reconciler | 25 apps per run, global | Enough for today's drift; not enough for thousands of apps |
UpdateOfferApiForAppJob | 60-second timeout, 3 tries. One failure retries the whole app | Large first syncs get cut short |
| Offer API HTTP | PATCH /v2/offers/{id} is 2000/min; GET, POST and DELETE on /v2/offers are 1000/min. Keyed by route and caller IP, so a dashboard instance shares one bucket | Bursts break it, not daily volume. The nightly metrics wave (00:00-02:00 UTC) peaks near the limit |
Realtime writes skip the HTTP limits entirely. Moving apps off batch removes their writes from the limiter.
Monitoring
Metrics come from new_dashboard with the prefix adgem_dashboard. (for example adgem_dashboard.eventsync.diff.drifted_offers) and from offer-api with offer_api.. Without the prefix a query returns no data rather than an error.
The EventSync monitors that are sound today:
| Monitor | Catches |
|---|---|
| Consumer unit stopped recycling | A hung consumer (it should restart every 9 minutes) |
| Queue backlog age | Messages older than an hour, for any reason |
| Nightly consume latency p99 | End-to-end latency, used to judge migration waves |
| Publishing has stopped | The dashboard stopped publishing. Known false positive in quiet hours; a replacement composite is ready but not yet imported |
Known gaps:
- a differ that runs but silently computes nothing;
- the DLQ monitor and the oversized-event monitor sit in No Data (the metrics don't report while idle), so they provide no real coverage;
- there is no SLO yet for drift or for time-to-inventory.
Tips for reading the metrics:
- Drift gauges are sparse, and Datadog averages them, so trust the direction of
drifted_offers, not its magnitude. For exact numbers, runeventsync:difffor the app. would_createrising whilewould_disablefalls, together, usually means the differ caught a supersede halfway (old row deleted, new one not yet written), not real drift.
Traps
Each of these has caused an incident or near miss.
- Adding a field under
goals, or changing any contract field in the formatter, supersedes every offer on every app. NeitherOFFER_API_IGNORED_FIELDSnor the differ can mask a field insidegoals. Plan such changes as a coordinated migration. - Adding a locale to
Language::nonEnglish()triggers a supersede storm on both paths, because translations are part of the contract. OFFER_API_IGNORED_FIELDSis the no-deploy switch for a field that diffs forever. It is comma-separated with no spaces (spaces make the key never match).- A per-app metric key must exist on both sides of the batch diff, on every offer, from the first slot. One side returning it while the other omits it is a diff on every offer forever; that was the 2026-08-25 storm. Adding one goes through the Add a Per-App Metric runbook, which uses the ignore list as a planned bridge.
- Never
horizon:clear --queue=defaultto stop a sync storm. It also drops EventSync publishes and user email. - Moving an app between paths is two edits: add it to one list and remove it from the other. Adding an app that was on neither list is a first-time enablement. That has a different blast radius, because the reconciler cold-starts its whole inventory.
- A Prism flag is not inventory. Before enabling Prism for an app, confirm it has live rows in the Offer API.
- DevCycle calls in
offer-apihave no timeout. Each flag check is a synchronous call to DevCycle's API, and nothing bounds it. Formatting a v1 offer response makes two, and the EventSyncCampaignMetricsChangedhandler makes one per app, so a slow DevCycle stalls API responses and the consumer alike. On 2026-09-29, a 15.7-second call onGET /v2/offersfailed a differ run. - An app can be
activeand soft-deleted at the same time. Laravel hides soft-deleted apps, but raw SQL (Metabase, Redshift) doesn't. When you build a list of apps to enroll, filter ondeleted_atas well asapp_status. - The Metabase mirror of the Offer API can freeze.
offer_api.offersin Metabase stopped updating for a week in September 2026 when a network rule was removed. Checkmax(updated_at)before trusting it. It also keeps superseded rows, so filter ondeleted_at IS NULL. For anything that matters, read the live Offer API or runeventsync:diff.
Where this is going
These goals are tracked in the Linear project EventSync: realtime sync for every app:
| # | Goal | Done when |
|---|---|---|
| 0 | EventSync is documented here | Someone new can enroll an app, verify its inventory and handle a sync incident from these docs |
| 1 | Retire batch sync | The batch schedule is empty and archived; sync-v2 --apps stays as break-glass |
| 2 | Realtime scales automatically | About 2,000 apps with no config change; a new app has inventory within an hour |
| 3 | Retire the hand-edited lists | Sync and Prism eligibility come from dashboard data, and Prism requires verified inventory |
| 4 | Correctness meets a measured SLO | The SLOs hold for the whole cohort, and the known correctness bugs are closed |
Scaling (goal 2) comes before retiring batch (goal 1), because the differ can't take the batch cohort yet. This page will change as that work lands.