Authenticating...
Skip to main content

EventSync Incident Response

When to use this​

Use this when:

  • an EventSync or offer sync monitor fires;
  • someone reports that an app's offers are missing, stale, wrongly paid, or still live after they should have stopped;
  • the Offer API is returning 429s to the dashboard.

How the system fits together is in Offer Sync Today. This page assumes you know the terms there: batch sync, realtime sync, the differ, the reconciler, contract fields.

The one thing to know first: realtime sync has a backstop. If the event path breaks, the differ still runs every 30 minutes and heals up to 25 apps per run over HTTP. So most EventSync failures mean slower and less correct, not an outage. Messages wait in SQS for up to 14 days. You usually have time to capture evidence before you fix.

Prerequisites​

  • AWS SSO access, region us-east-2:

    • the Offer API account (010438502987), environments offer-api-green-production (web) and offer-api-green-production-worker (worker, where the consumers run). Commands below use the profile name adgem-offer-api.
    • the AdGem dashboard account, environment adgem-dashboard-production. Profile name adgem-legacy.
  • ssm:SendCommand on those instances. Instance ids change on every deploy, so look them up each time:

    aws elasticbeanstalk describe-environment-resources \
    --environment-name offer-api-green-production-worker \
    --profile adgem-offer-api --region us-east-2 \
    --query 'EnvironmentResources.Instances[].Id' --output text
  • Datadog, and edit access to the adgem project in DevCycle.

Triage: which alert fired?​

MonitorWhat it meansGo to
EventSync: consumer unit stopped recycling (wedged)A consumer process is alive but stuck. Its queue is fillingConsumer wedged or dead
EventSync: queue backlog ageMessages older than 15 min (warn) or 1 h (crit) on a queue. Nothing is draining it, for any reasonConsumer wedged or dead
EventSync: consumer receiving but not processingA handler is throwing, or the consumer crashes mid-messageHandler failures and the DLQ
EventSync: DLQ has messagesAn event failed every retry and was set asideHandler failures and the DLQ
EventSync: publishing has stopped (dashboard side)The dashboard published nothing for 4 hoursPublishing stopped
EventSync: nightly consume latency p99 (migration gate)The nightly burst drains too slowly. Not an outageSlow drain
EventSync: reconciler shedding apps at per-run capMore apps drifted than the reconciler heals per runDrift and reconciler shedding
Cron: hourly-band scheduled job stopped running (eventsync:diff)The differ hasn't completed in 3 hoursDiffer not running
Offer API 429s / dashboard "Offer API request failures"The HTTP paths (batch, reconciler) exceed the Offer API rate limitSync storms

Two monitors don't give real coverage today, so don't read their state as evidence: "oversized EventBridge event refused" (no data for over 30 days), and "DLQ has messages", whose series lapses while the DLQs are idle.

Consumer wedged or dead​

There are two consumers, eventsync-consumer@update and eventsync-consumer@supersession, systemd units on the Offer API worker. A healthy unit exits and restarts every 9 minutes (--max-time=540).

  1. Confirm with SQS. Visible messages with zero in flight means nothing is receiving:

    for q in offer-api-update-events-production offer-api-supersession-events-production; do
    url=$(aws sqs get-queue-url --queue-name "$q" --profile adgem-offer-api --region us-east-2 --query QueueUrl --output text)
    aws sqs get-queue-attributes --queue-url "$url" \
    --attribute-names ApproximateNumberOfMessages ApproximateNumberOfMessagesNotVisible \
    --profile adgem-offer-api --region us-east-2 --output text
    done
  2. Check the units on the worker (through SSM, AWS-RunShellScript):

    systemctl status eventsync-consumer@update eventsync-consumer@supersession --no-pager
    df -h /
    • An uptime of more than about 10 minutes means the unit is wedged: it never reached its restart deadline.
    • In Datadog, systemd.unit.active reads a flat 1.0 for a wedged unit. A healthy one dips every 9 minutes. A flat line is the failure, not health.
    • A full root disk kills the app while the instance looks healthy (2026-08-13).
  3. Capture evidence before restarting, if you have time. A restart destroys it. With PID from systemctl status:

    sudo ss -tanpo | grep "pid=$PID"
    sudo cat /proc/$PID/stack
    sudo ls -l --time-style=full-iso /proc/$PID/fd
  4. Restart:

    sudo systemctl restart eventsync-consumer@update

    It takes about 90 seconds. A stuck PHP process ignores SIGTERM, so systemd waits out its stop timeout and then kills it. That's expected.

  5. Confirm recovery: in-flight goes above zero within a couple of minutes, and the visible count falls. A backlog of about 3,300 messages drained in about 10 minutes on 2026-08-15.

Nothing is lost in either case. Once the consumer is back, the differ clears any drift left behind at 25 apps per run.

Handler failures and the DLQ​

  1. offer_api.eventsync.consume.errors by detail_type names the failing event type.
  2. Search service:offer-api error logs for that type.
  3. Before purging a DLQ, save the messages. They're the only copy. Read them in the SQS console, or with aws sqs receive-message on offer-api-update-events-dlq-production / offer-api-supersession-events-dlq-production.
  4. After the handler is fixed, redrive the DLQ back to its source queue (SQS console, Start DLQ redrive). Handlers are idempotent, so replaying is safe. If you can't replay, the differ heals the affected apps anyway; list them with eventsync:diff (see below).

Publishing stopped​

Work down the checklist in the monitor's message. In short:

  1. Was the cohort emptied on purpose? adgem_dashboard.eventsync.cohort_size, and the DevCycle audit log for real-time-offer-sync-apps. Turning that feature off does not empty the cohort: the code keeps the last good list.
  2. Is the differ running? See Differ not running.
  3. Is the dashboard failing to reach EventBridge? aws.events.invocations and aws.events.failed_invocations for the bus adgem-offer-sync-production.

The dashboard ships no info-level logs, so quiet logs prove nothing.

Slow drain​

The consumer's cost per message grows with the realtime cohort. The nightly CampaignMetricsChanged burst (00:00-01:00 UTC) is where it shows.

  • A warning means hold further migration waves. It isn't an incident.
  • Check cohort_size first: if the cohort just grew, that's the answer.
  • Then check offer_api.eventsync.consume.duration_ms for campaignmetricschanged. A rise with a flat message count means the handler is doing more work per message.
  • A slow DevCycle slows the drain too: the CampaignMetricsChanged handler checks a feature flag for each app, with no timeout.
  • A slow drain with a healthy unit points to a degraded worker instance. Compare with the previous nights.

Drift and reconciler shedding​

The reconciler heals at most 25 apps per run, across all shards. Apps over the cap wait for a later run, so shedding means healing is delayed, not lost.

  1. Is the event path healthy? Sustained shedding usually means the consumers are down. Check that first.

  2. What kind of drift? adgem_dashboard.eventsync.diff.drifted_offers by action, and drifted_fields by field. Query with .as_count(); without it Datadog averages the gauge into fractions.

    • would_disable: stale offers still being served. The most serious.
    • would_supersede: offers served with a stale contract field, for example an old payout.
    • A one-run burst across many apps after a bulk campaign edit usually heals within one or two runs. That's normal.
  3. Get exact numbers for an app from the command, not the graphs:

    php artisan eventsync:diff 1234 5678 --show-noop

    Run it on a dashboard instance through SSM. It's read-only without --reconcile, takes positional app ids, and exits 0 even when it finds drift.

  4. Heal specific apps now, ahead of the cap: offer-api:sync-v2 --apps=1234,5678. See step 2 of Enroll an App.

Raising the cap (OFFER_API_RECONCILE_MAX_APPS_PER_RUN) is rarely the fix. The reconciler writes over HTTP into the Offer API, which is usually already struggling when the cap is under pressure.

Differ not running​

  1. adgem_dashboard.scheduled_task_failure{command:eventsync:diff}:
    • failures present: the command runs and errors. Read its logs.
    • failures absent: it isn't running at all. Suspect the scheduler, or a held withoutOverlapping lock (25 minutes).
  2. Suspect a slow DevCycle. Feature flag checks in the Offer API are synchronous calls to DevCycle's API with no timeout, so a slow DevCycle stalls whatever is waiting on it. The differ reads every app's offers from the Offer API, so it stalls too. On 2026-09-29, a 15.7-second DevCycle call on GET /v2/offers failed a differ run. Check DevCycle's status page.
  3. While the differ is down, realtime apps have no backstop, but the event path keeps working.

Sync storms​

Symptoms: 429s from PATCH /v2/offers/{id} (limit 2,000 a minute), UpdateOfferApiForAppJob failures, a spike in would_supersede, and a Horizon default queue backing up on the dashboard.

Usual causes, from past incidents:

  • A field that diffs forever. A key inside offer_metrics that one side of the batch diff carries and the other does not: the Offer API returned rpc_d7_per_app on every offer while the dashboard sent it only for some, so every slot re-patched every offer (2026-08-25, PUB-660). The mirror case, the dashboard sending a key the Offer API does not know, is not a loop but a 422 on every metrics PATCH. The procedure that avoids both is Add a Per-App Metric.
  • A bulk campaign edit burst. A campaign's offers are mirrored on every app that runs it, so one edit fans out to up to about 190 writes. The burst rate matters, not the daily volume. It was first blamed for 08-25 and was not the trigger, but it does amplify any loop that is already running.
  • A contract change. A new field under goals, or a new locale in Language::nonEnglish(), supersedes every offer on every app.

Levers, from least to most disruptive:

  1. OFFER_API_IGNORED_FIELDS: excludes fields from the dashboard's diff, with no deploy. It's an Elastic Beanstalk environment property on adgem-dashboard-production. On 2026-09-29 it was non_linear,is_purchase_goal,rpc_d7_per_app; PEX-683 removes the rpc_d7_per_app entry once the per-app key is sent on every offer. Adding a new per-app metric uses it as a planned bridge: see Add a Per-App Metric.
    • Comma-separated, no spaces: a space makes the key never match.
    • Changing it restarts the environment, about 10 minutes. Afterwards, check the value took with aws elasticbeanstalk describe-configuration-settings.
    • It cannot mask a field inside goals.
    • It also stops real changes to that field from syncing, so remove it again once the code is fixed.
  2. Turn off the reconciler with the DevCycle flag eventsync-reconcile-enabled. The differ keeps reporting but stops dispatching HTTP heals. Use it when the reconciler is adding load to an Offer API that is already failing.
  3. Batch apps: there's no pause switch. Removing apps from offer-api-v2-sync-schedule stops their future runs, but that flag has other editors and a bad edit silently drops apps. Treat it as a last resort, and write down exactly what you removed.
Never clear the dashboard's default queue

horizon:clear --queue=default drops everything on that queue: EventSync publishes, transactional email, and Cognito registration, along with the sync jobs. There's no per-job purge.

Past incidents​

Date (UTC)What happenedLesson
2026-08-13 to 08-15Worker root disk filled with unrotated logs. The app died, the instance stayed "healthy", and both consumers were down 57 hoursCheck the disk. The reconciler carried the whole load and shed apps for the entire outage
2026-08-25 to 08-26rpc_d7_per_app shipped on both sides 21 seconds apart: the Offer API returned it on every offer, the dashboard sent it only for some, and the batch diff nulled it on every offer every slot. About 51,000 PATCH per 30 minutes for 21 hours, 0.02% of them writing anything; the 429s and whole-app retries were a consequence, not the trigger (PUB-660)A key on one side of the diff only never converges. Stopped with OFFER_API_IGNORED_FIELDS within minutes of diagnosis; fixed for good by sending the key on every offer (PEX-571, PEX-683). Procedure for the next metric: Add a Per-App Metric
2026-09-03, 09-21French, then Italian, added to Language::nonEnglish() superseded offers across every app on both pathsTranslations are contract fields. Plan locale additions
2026-09-17 to 09-18eventsync-consumer@update blocked on a dead SQS connection for about 21 hours. Every existing monitor stayed greenFixed with an HTTP timeout on the SQS client. The wedge monitor now catches it within 15 minutes
Reported 2026-09-24Concurrent creates had left two live offers for 157 app and campaign pairs, some dating from 2024Fixed 09-29 with a per-pair lock and a unique index

Known gaps​

  • A differ that runs to completion but computes nothing isn't detected.
  • There's no SLO yet for drift or for time-to-inventory. It's being defined in the EventSync project.
  • The routing of each event type to the update or supersession queue is set in infrastructure code and isn't documented here yet. campaignmetricschanged goes to update.