0079: Heuristic Rule Layer for Conversion Auto-Approval
STATUS
Accepted
Records the target architecture for a detection capability that does not exist yet. Nothing in the DECISION section is built. The measurements quoted in CONTEXT are from production as of 2026-08-31.
CONTEXT
The problem
Conversion auto-approval is a rules engine, not a model. Every pending AdGem conversion is scored against a versioned list of exclusion rules; anything matching nothing is approved in TUNE automatically, anything matching is held for a human. Today that list is 77 rows across seven matchable dimensions — offer, app, app base, publisher, device brand, device model, OS version — plus payout thresholds and NULL guards.
Every one of those rules answers the same question: is this attribute on a list? Which means a rule only exists after a person noticed a bad offer or a bad device and added it. The engine is reactive by construction, and it cannot express anything about a conversion's behaviour — how fast the goal completed, how many conversions that player fired in ten minutes, whether the click came from a country the offer does not target.
Those behavioural signals already exist. timing_fraud_signals,
individual_player_fraud_signals, location_fraud_signals and skipped_goals_fraud_signals were
built as ML features for the fraud report. They are not wired into approvals.
What a heuristic is
A computed, thresholded predicate over the conversion and its history, emitting a named rule hit. It differs from an exclusion rule in that its input is derived rather than looked up, and from an ML score in that it is a single explainable condition rather than a learned weighting.
Decision drivers
- It has to gate, not merely observe. A signal that arrives after the approval has been sent to TUNE cannot prevent anything.
- It must not slow the approval loop significantly.
autoapprove_conversionsis dataset-triggered onevents_streamroughly every 10–15 minutes,max_active_runs=1, re-scoring the same trailing 24-hour window around 144 times a day. The working budget is single-digit minutes. Some cost is unavoidable — the decision below adds a task to the DAG — so the bar is that the loop stays comfortably inside its budget, not that nothing is added. - It has to be measurable per rule. "This rule catches fraud" is not a claim we can act on without knowing what it costs and what it adds over the rules we already run.
- It has to be explainable, in the mechanism that already exists. Concretely: the rule's name
lands in
reason, which the fraud report surfaces asnot_autoapproved_reason. A reviewer sees why a conversion was held in the column they already read, with no new surface to learn. This is a requirement, not a preference — an unexplainable hold cannot be adjudicated by the person the hold is asking to adjudicate it. - It has to fail safe. Exclusionary config fails open — a partial load silently widens approval rather than erroring (ADR-0073, Incident 203). Anything added here inherits that hazard.
- It must use only information available at decision time. The same point-in-time contract the ML feature work is held to.
Considered options — where the layer executes
Option A: a standalone heuristics service. Own deployment, own scaling, calls out to the warehouse. Maximum isolation from the approval path. Rejected: it duplicates the rule-evaluation and hold-logging machinery that already exists, introduces a network dependency into a path that has none, and gives us two systems that can disagree about one conversion.
Option B: a separate Airflow DAG, scheduled on the evaluation's dataset. The seam already
exists — refresh_autoapproval_evaluation emits
redshift://adaction_analysis/conversion_review/conversion_autoapproval_evaluation as an auto-outlet
and nothing consumes it. Rejected for gating: approve_conversions is downstream of that
evaluation and sends IDs to TUNE immediately, so a DAG triggered on the outlet runs after the
approval has already happened. It could flag or reverse. It could not hold.
Option C: a fourth arm of the blocks union inside conversion_autoapproval_evaluation.
blocks is already a heterogeneous union — the NULL guards and the payout thresholds are non-lookup
rules emitting literal names, so a heuristic arm would be a fourth producer of the same
(tune_event_id, rule_name) shape rather than a new concept. Rejected in favour of E, for reasons
the comparison below cannot see.
Option D: extend the rules CSV to carry thresholds. Rejected for now. A rule row is one
dimension and one value; it has no column for a threshold, a window, or a scope. The
attapoll_risk tier is the scar tissue from working around exactly that gap — it required a new
tier in three places plus a hardcoded app ID in SQL. Thresholds stay in code until they are stable
enough that the argument shifts from correctness to operations.
Option E: a separate heuristics model, combined with the exclusion rules downstream. Chosen.
The heuristics model emits (tune_event_id, rule_name) for the pending snapshot, the evaluation
model emits exclusion-rule hits unchanged, and a third model combines them into one outcome. Every
property C has is preserved — it still gates, it still writes into one reason, it still logs holds
— because the combined model still runs upstream of the approver.
Joining the heuristics model into the evaluation's blocks union would be C wearing a different
hat: two unrelated computations in one model, with the evaluation still waiting on the heuristics.
The combining belongs downstream of both. See decision 2.
Considered options — how it rolls out
Shadow only. Compute the flag, change nothing. Zero risk, zero review cost. Rejected as an endpoint: the conversions a heuristic would hold are auto-approved and never reviewed, so it measures how often a rule fires and nothing about whether it is right.
Hold-only from the start. Every flagged conversion is held, lands on the fraud report, and comes back with a verdict. Produces labels, but it makes the product decision in order to discover whether the decision was justified.
Staged. Chosen. See decision 5.
Comparison
| Gates? | Loop cost | Needs new infra | Two systems can disagree | |
|---|---|---|---|---|
| A: standalone service | Yes | External call | Yes | Yes |
| B: separate DAG | No | None | No | Yes |
C: fourth blocks arm | Yes | Predicate only | No | No |
| D: extend the CSV | Yes | Predicate only | Loader + DDL | No |
| E: separate model, combined downstream | Yes | Predicate only, + one task | No | No |
These criteria do not separate C from E, which is worth stating rather than hiding: they score identically on all four. The choice comes from two things the table does not measure.
Consistency. Decision 3 holds that the expensive, slow-moving part of a heuristic is a feature and belongs in its own model. E applies that same boundary to the predicate. There is no principled reason for the predicate to sit on the far side of a line drawn one step earlier.
Fail-safe by construction. Under E, abstaining is the default rather than a behaviour that has to be written: a model with stale inputs emits zero rows, the join finds nothing, and the evaluation carries on. Under C the same guarantee is a defensive check living inside the file that decides every auto-approval.
There is also precedent. recent_conversion_events_stream_data is a separate model for exactly this
reason, in its own words: its own DAG task, which keeps that cost and its failure signal separate
from the evaluation step. E is the pattern this flow already uses one step upstream; C would be the
deviation. The cost of E is one additional DAG task per run, roughly 144 times a day.
DECISION
1. Heuristics are a first-class rule class, not a stopgap until ML
They are explainable, changeable in hours, and cheap to retire. The ML path ranks residual risk they miss. Neither replaces the other. The full argument for maintaining both paths belongs to its own ADR (PEX-486); this one states the position it depends on rather than making the case.
2. Three models: hard rules, heuristics, and a combined outcome
The two rule classes compute in parallel, and a third model combines them.
| Model | Produces | Grain |
|---|---|---|
conversion_autoapproval_evaluation | Exclusion-rule hits, unchanged | One row per pending conversion |
conversion_heuristic_evaluation | Heuristic rule hits | One row per (conversion, heuristic hit) |
conversion_approval_decision | The conversion's outcome and the evidence behind it | One row per pending conversion |
Naming, so this is settled once rather than in every review: the heuristics model takes the
conversion_*_evaluation shape of the model it runs beside, because it does the same job on the
other rule class. The third is a decision rather than an evaluation because that is the
distinction decision 7 rests on — evaluations emit hits, the decision model turns hits into a
verdict. Bikeshed them here if you want them different; downstream measurement SQL will point at
these names, so they should stop moving.
Folding heuristics into the evaluation model as a fourth arm of its blocks union is rejected.
The two models compute different things — one looks values up in a list, the other
thresholds a computed signal — and joining one into the other makes the evaluation unable to start
until the heuristics have finished, serialising two things that have no reason to be sequential.
The combined model carries a source column, so the layer that produced a piece of evidence is
a field rather than something inferred from a rule's name. hard_rule and heuristic today;
ml_score when Path B arrives, with no redesign. Nothing downstream should have to parse reason
to learn where a hit came from.
Heuristics evaluate every pending conversion, not a subset. It would be tempting to run them only over what the exclusion rules approve, since that is where leakage lives. That would be a mistake: scoring the whole snapshot costs nothing extra — it is the same 24-hour table — and it lets the combined model classify each conversion as hard-rule-only, heuristic-only, or both. That classification is the incremental-value measurement in decision 5. Restricting the input throws half of it away.
What the combined model preserves:
autoapprove, computed from evidence of either kind;reason, still the alphabetically-orderedLISTAGG(rule_name, ', '), so sole-reason attribution stays an equality test on one column;- permanent hold logging, and a visible reason on the fraud report.
The constraint on reason is about rule hits specifically, and the distinction matters for Path
B. An ML risk score is not a rule hit — it is a continuous value carried with its model and feature
versions, consumed by the decisioning layer. It does not compete for reason, and nothing here
forecloses it.
One practical note: reason is varchar(512) and the evaluation model wraps its LISTAGG in
cast(... as varchar(512)), so overflow truncates silently rather than erroring. That is the
worse of the two failure modes for this design — a silently shortened reason breaks sole-reason
attribution and every measurement built on it, with nothing failing to signal it. Widen the column
rather than designing around it; observed max today is 81 characters, so there is room to do it
deliberately. The 33 existing unit tests that assert
reason verbatim are the real constraint on its format, and they should keep passing unchanged.
3. Peer baselines are their own model, incremental, on the same cadence
The heuristics model may reference the trailing-24h pending snapshot and a precomputed baseline. It
may not compute peer statistics inline. timing_fraud_signals and skipped_goals_fraud_signals_v2
build an unbounded per-subject peer join over a 180-day window before a 5,000-row cap trims it, and
that does not fit a loop running every ten minutes.
Naming the thing, because "precomputed upstream" is a description rather than a specification:
| Model | offer_goal_timing_baseline (and a sibling per signal family that needs one) |
| Grain | One row per (offer_id, goal_id, conversion_date) |
| Columns | peer_count, sum_x, sum_x2 per day; mean and stddev derived over the trailing window; window_start, window_end, computed_at |
| Population | The (offer_id, goal_id) pairs the pending snapshot actually contains over a trailing window — not every pair in events_stream |
| Materialization | Incremental on conversion_date |
| Schedule | The same cadence as the model that reads it |
Incremental, and on the same cadence as its reader. A daily full rebuild read from inside the ten-minute loop would buy little and cost a staleness problem needing its own guard, its own alarm and its own failure mode.
The reason it can run on one cadence is the shape of the cost. What makes the peer statistics
expensive today is the per-subject fan-out — roughly 5,000 peer rows materialised per subject before a cap
trims them. A (offer_id, goal_id, day) aggregate does not have that shape at all: new conversions
fold into existing buckets, so the marginal cost per run is a few thousand rows rather than millions.
Cheap enough to refresh on the same cadence, which removes the cadence gap instead of managing it.
Storing daily moments rather than a single mean and stddev is what makes it genuinely incremental while keeping a trailing window: the window is derived by differencing the cumulative values at its two ends, so a new day is an insert rather than a rebuild and an expiring day falls out by arithmetic. That is the same technique PEX-483 is considering as its Option B for the upstream fan-out, which is worth noticing — one fix, two consumers.
The split that still matters is between the baseline, which is per offer/goal and an aggregate, and the per-conversion part, which is a z-score against that lookup. The second is arithmetic and belongs inline. Nothing per-conversion is precomputed; nothing with a per-subject fan-out is computed inline.
The population is defined, not implied. Tens of thousands of conversions arrive daily, not
every goal reaches auto-approval, and not every offer is active — so building a baseline for every
pair in
events_stream would pay for rows nothing can ever read. The population is the distinct
(offer_id, goal_id) pairs present in the pending snapshot over a trailing 30 days, as a dbt
var. That is self-maintaining: a goal that stops appearing ages out, and a new one appears the first
time auto-approval sees it.
Thirty, because the curve is flat after it. The trade-off runs both ways — too short and an offer
with intermittent traffic loses a baseline it could have kept, abstaining when it need not; too long
and rows are maintained that nothing reads. Measured against the current pending snapshot, of the
1,901 distinct (offer_id, goal_id) pairs in it:
| Window reaching back | Pairs already seen | Share |
|---|---|---|
| 14 days | 1,820 | 95.7% |
| 30 days | 1,845 | 97.1% |
| 60 days | 1,847 | 97.2% |
| 180 days | 1,848 | 97.2% |
Going from 30 days to 180 recovers three pairs out of 1,901. The remaining 2.8% have no prior
history at any window length — they are genuinely new offer/goal pairs, and they abstain on
min_peer_count below, which is the intended behaviour rather than a gap the window could close.
This is not the 180-day peer window, and the two are easy to conflate. They answer different questions: the population window decides which pairs get a baseline at all, the peer window decides how far back conversions are aggregated into one. A pair can be in the population on the strength of one conversion last week and still have 180 days of peers behind it.
Thin data abstains on its own terms, not on staleness. A new offer/goal has few peers, and
cutting it off for staleness is the wrong instrument — it makes an offer with no baseline
indistinguishable from a table that failed to build. So the guard is
peer_count >= min_peer_count, a dbt var: below it the heuristic emits nothing for that offer/goal
and says so as insufficient support. A z-score over a handful of peers is not a weak signal, it is
noise with a denominator, and the honest answer for a brand-new offer is that this rule cannot speak
to it yet.
Freshness is still checked, as a backstop. The heuristics model joins the baseline with a bound
on computed_at, and a baseline older than the bound produces no row — the abstain in decision 6.
With the baseline on the same cadence this should never fire in normal operation, which is the
point: it now catches a broken build rather than papering over a design gap.
Abstaining silently is not acceptable on its own. A heuristic that quietly stops firing looks identical to one that found nothing, and it would be worse than useless if the review team came to rely on it. The freshness guard is the safe behaviour; a separate alarm is what makes it visible.
The alarm has to be a check that fails, not a callback waiting for one.
autoapprove_conversions carries an on_failure_callback to the ML team channel, but it is a
DAG-level hook that fires on run failure, and a stale baseline fails nothing — the join finds no
row, dbt succeeds, the run goes green. So staleness needs a task of its own that raises, in the
shape test_autoapproval_rules already has in that DAG for the same class of problem: something
that would otherwise pass green having silently done nothing. The existing hook then carries it.
Downstream of the approval task rather than upstream of the evaluation, so a stale baseline alarms
without stopping the cycle — abstaining is the safe degradation, not a full stop.
This relaxes the point-in-time contract, and that is worth stating plainly rather than leaving to be found. Driver 6 claims the same contract the ML features are held to. A precomputed group baseline is not that contract:
- Today
subject_peer_statsaggregatesgroup by subject_tune_event_id— every subject gets its own window, ending strictly before it. That is what PEX-426 established. - A
(offer_id, goal_id)baseline is one snapshot shared by every conversion on that offer and goal until it is rebuilt. Two conversions hours apart get the same denominator where today they get different ones.
The looseness is probably immaterial at 5,000 peers over 180 days, but "probably" is doing real work in that sentence and it should be measured rather than assumed.
Measured how matters. Average disagreement across the distribution is not the number to look at: a heuristic only fires in the tail, so two baselines can agree on 99% of conversions and still disagree on a third of the ones that trip the rule. The check is the disagreement rate conditional on either z-score exceeding the threshold — the rate at which the group baseline and the subject-relative one would produce different hold decisions at the threshold being shipped. Same window, same threshold, reported as a rate.
It also means decision 4's shared macro shares the arithmetic, not the baseline. The ML feature path keeps subject-relative peers; the heuristic path scores against the group snapshot. The formula is one definition; the denominator is not. If the comparison above shows the two diverge materially, that is a finding about the heuristic path, not a reason to relax what PEX-426 fixed.
4. A macro computes the indicator; a model applies the threshold
The two are separate jobs. A threshold is not part of an indicator's definition, so a macro does not take one as an argument.
The macro computes the indicator — the z-score, the count, the geo match — and nothing else. It is called over the reviewed population by the fraud-signal model, where the indicator is an ML feature, and over the pending population by the heuristics model, where it is about to be thresholded. Same definition, two callers — and eventually three: when Path B scores pending conversions it needs the same indicator over the same population, so a served ML model becomes a third caller of the identical macro rather than a reimplementation of it. That is the main reason this boundary is worth drawing now rather than at the point of serving.
The heuristics model applies the threshold, because thresholding is what turns an indicator into a rule and only the production gate needs it. The training path wants the raw indicator; a threshold baked into the shared macro would be dead weight on one side and an invisible coupling on the other.
Thresholds live as dbt vars read by that model, which keeps a threshold change a reviewable diff — load-bearing while precision is being measured at a threshold, since a silent mid-window change invalidates the measurement.
Why the macro matters at all: the five existing signal models all drive off
manually_reviewed_events_stream, an inner join to the manual review report, so a pending
conversion has no row and none of them can score what auto-approval is deciding on. The logic is
reusable; the population is not. A macro makes re-pointing a new caller rather than a fork that
drifts. Per decision 3, what the macro shares is the arithmetic — the baseline behind it differs by
population.
5. Rollout is staged, and each stage answers a different question
| Stage | What it does | What it measures | Product decision |
|---|---|---|---|
| 0. Offline | Join the existing signal output to historical verdicts | Precision on the overlap population; overlap with current rules | None |
| 1. Shadow | Predicate on the pending population, own cadence, own record | Volume, stability, and a flag reviewers can see and dispute | None |
| 2. Sampled back-review | A uniform random sample of conversions the rule flagged and the approver took, handed to review as a one-off batch | Precision on the incremental population | None — but the stage is required, not optional |
| 3. Hold-only | The arm clears autoapprove | Sole-reason volume and precision, ongoing | Yes |
Stage 0 costs a query and can end a bad rule before any production code exists. Stage 2 is the one that is easy to skip and must not be: stages 0 and 1 only see conversions the existing rules already hold, and the entire justification for a heuristic is the conversions they do not. A rule validated on the overlap population and shipped against the incremental one has not been validated.
Stage 2 is a gate, and it is small. Left unsized it reads as a large ask of the review team, and therefore as the stage most likely to be dropped under time pressure. Sized against production it is not large:
| Sample | Interval | Keep/drop gate vs the 0.2097 baseline | Review time at ~1,050/2h |
|---|---|---|---|
| n=100 | ±9pp | 29 confirmed fraud out of 100 | ~11 minutes |
| n=200 | ±6.4pp | proportionally | ~22 minutes |
Back-review is likely slower per item than the normal queue, because there is no rule name to anchor on. Even at three times the rate it is under an hour per rule. That is the whole cost of knowing whether a heuristic works on the population it exists to catch, so no rule reaches stage 3 without it.
Two conditions on how the sample is drawn, because both change what the number means:
- It runs at the threshold being shipped. A precision estimate does not carry across a threshold change, so a sample drawn at one threshold cannot license a different one.
- It is uniform random over the flagged-and-approved set, not top-N by score. Top-N measures the rule at its most confident, which is precisely not the population the threshold will admit.
What stage 3 can honestly measure today is narrower than "sole-reason volume and precision."
reason carries the latest evaluation, and a pending conversion is re-scored on every cycle — in
practice many times, see the follow-up on rule-level attribution. For a time-invariant predicate
that is harmless: geo targeting and completion time do not change while a conversion waits, so the
last evaluation is every evaluation. For a time-varying one it is not: a velocity count is a
different number an hour later, so a rule that fired early can leave no trace by the time the record
settles.
The consequence is a sequencing constraint rather than a redesign. Start with a time-invariant
predicate — click_in_offer_country is the obvious first rule and is already a macro — so a
heuristic can ship and be measured before per-run attribution exists. Time-varying families wait for
it.
No heuristic gets rejection authority at any stage.
6. Stale features abstain; a failed model stops the cycle
These are two different situations and they get two different answers. Conflating them is how a fail-safe ends up asking for both "carry on as before" and "stop everything" in the same breath.
| Situation | Behaviour | Mechanism |
|---|---|---|
| A baseline has too few peers | Abstain — no hits for that offer/goal, recorded as insufficient support | The min_peer_count guard in decision 3 |
| A baseline is stale or missing | Abstain — no hits emitted for that heuristic | The computed_at bound in decision 3 excludes it, so the join finds nothing and the model emits no row |
| The heuristics model or its task fails | Nothing is approved this cycle | Task failure leaves the combined model upstream_failed, as with any upstream task |
Abstaining is the safe default because it degrades to exactly today's behaviour: the existing rules decide, as they do now. Failing loudly is the safe answer to a broken model, because the alternative is approving on output nobody verified.
Safe here means safe to run, not safe to depend on. Abstaining protects the approval decision; it does nothing for anyone who has started assuming the heuristic is watching. That risk grows the moment the layer is useful enough that exclusion rules stop being maintained as closely — so the alarm in decision 3 is a condition of relying on this layer, not a nice-to-have that follows it.
Decision 2 is what makes the first row cheap rather than defensive. A separate model with a stale baseline emits nothing, and emitting nothing is already the correct behaviour — there is no staleness branch to write inside the evaluation, and no risk of getting it wrong there.
7. The heuristic emits a rule hit; deciding what to do with it belongs elsewhere
Three components, named once and used consistently for the rest of this ADR. The two evaluation
models (conversion_autoapproval_evaluation and conversion_heuristic_evaluation) each emit rule
hits and nothing else. The decision model (conversion_approval_decision) computes the verdict
from every hit of either kind. The approver (ConversionApprover) acts on that verdict against
TUNE.
A heuristic contributes a name to reason and nothing else. Today the only outcome that name can
produce is hold-or-not, and that is the combined model's to decide, not the heuristic's. Any richer
outcome is a decisioning-layer concern — see below for what a heuristic is allowed to do, and the
ADR covering both paths for the cross-path version once an ML risk score is also in play.
Establishing the separation now is cheaper than retrofitting it, and Path B needs the same one: a model that emits a score rather than a decision.
How thresholds are shaped
One threshold per indicator, and each indicator emits its own rule_name. No heuristic depends
on another, and no hold requires two indicators to fire together.
It is a deliberate limit, and worth being precise about why — because one tempting reason for it is
wrong. A combination is not unmeasurable by the sole-reason method: sole-reason attribution
constrains naming, not arity, so a combination emitting its own rule_name is measured by exactly
the same equality test on reason as any single-indicator rule. What is true is narrower — you
cannot measure a combination's individual components from inside it.
The reasons that do hold are smaller: one threshold per indicator is what the threshold work in the follow-ups is scoped to produce, independent predicates union together where combined ones branch, and a per-indicator catalog is what makes the incremental-value question answerable rule by rule.
The limit costs something measurable. In the actionable band, single-rule holds and two-rule holds are not close:
| Rules fired | Held | Reviewed | Precision |
|---|---|---|---|
| 1 | 3,057 | 1,674 | 0.2097 |
| 2 | 792 | 495 | 0.7697 |
Conversions tripping two rules are roughly 3.7x more precise than those tripping one. That does not make combinations v1 work, and the ordering is not arbitrary: the information is recoverable in one direction only.
Ship two indicators independently and a combination can later be measured against both of their sole-reason baselines — its incremental value over each is a subtraction. Ship the combination first and there is no baseline for either component to subtract from, and per the correction above you cannot recover one from inside the combination. So combination-first forecloses a measurement that component-first preserves. That is the reason, rather than convenience or v1 scope.
Combining weak signals into a single verdict is also what Path B is for. But the table above means the independence limit is a sequencing decision with a known price rather than a free architectural boundary, and it should be revisited on its own schedule rather than left to Path B by default.
The condition for revisiting it is checkable, not a feeling. A combination becomes measurable once each of its components has a sole-reason precision measured at the threshold it actually shipped with, from stage 3 over at least one full review cycle. Before that there is nothing to subtract from; after it, the combination's incremental value is arithmetic. That is the same criterion as the paragraph above, stated as a gate.
Heuristics carry a payout floor. The argument against one — that a second knob moving during a measurement window makes the result harder to read — is methodologically tidy and practically wrong.
The report is already a targeted population: most traffic under $40 is never reviewed at all. A heuristic firing below the floor would not be adding a variable to a population under review — it would be adding a population nobody looks at. A thousand fifty-cent conversions represent at most a few hundred dollars of exposure while displacing attention from the band where the money is.
Note the argument is opportunity cost, not capacity. Review throughput is not currently the binding constraint — the review team reports finishing a day's queue in about two hours. The floor is about what reviewer attention is worth spending on, which is why "we have capacity" does not answer it.
A floor is a second condition, so it is worth being precise about what the independence limit above
actually forbids. Combining detection signals — two indicators together forming a verdict
neither reaches alone — is what Path B is for, and no heuristic does it. Scoping — where a rule
applies at all — is a different thing, and payout is scope rather than evidence. The exclusion rules
already work this way: a medium_risk rule is one value in one dimension, gated to a payout band.
The independence limit is revisitable, on the condition stated above — once each component rule has a sole-reason precision measured at a shipped threshold. The floor is not: it is a property of what the review process is for, rather than of how well tuned a rule is.
What the decisioning layer is, in Path A terms
For heuristics it is not hypothetical and not future work. It is conversion_approval_decision
producing a combined verdict from every rule hit of either kind, and ConversionApprover acting on
it. A heuristic is one more input to that decision. The cross-path
version — deterministic rules, plus heuristics, plus an ML risk score, plus business policy — is a
different question and belongs with the ADR that argues for both paths.
What a heuristic is allowed to do
Two actions exist today, and they sit at opposite ends of a boundary worth naming:
| Action | What decides it today | Undoable by a human? |
|---|---|---|
| Auto-approve | Nothing fired | No — the player has been paid |
| Hold for review | Any rule hit | Yes — a reviewer approves it |
| Auto-reject | Not admitted | No |
The useful ordering is by reversibility, not severity. A hold made in error costs a player some waiting and a reviewer some minutes, and a human undoes it. Auto-approve and auto-reject are the two ends where a mistake is not recoverable by review — one pays a fraudster, the other refuses a legitimate player, and neither gets a second look by default.
That boundary, rather than caution, is why auto-reject stays out of scope. It gives a rule that outlives this ADR: a detection source may move a conversion up to the highest action its evidence justifies, but never across the reversibility boundary without a human in the path.
A heuristic can move a conversion from auto-approve to hold, and nothing else.
Richer actions — ranking within the review queue, routing to a different reviewer — are plausible and deliberately not named here: naming actions we have neither defined nor agreed reads as decided when it is not. What those actions should be, whether they are distinct from each other at all, and what each is worth belong to the conversion review team, who own the queue and are the end users of both detection paths. Tracked separately; see Follow-Ups.
Flow
CONSEQUENCES
Positive
- Detection stops being purely reactive: a rule can express behaviour, not just membership of a list.
- The signal work already built for ML earns a second consumer, with one definition shared by both.
- Every heuristic is measurable per rule from day one, because sole-reason attribution comes free
with the
reasonstring. - Reviewers see the reason in the column they already read; no new surface to learn.
- The staged rollout means no product decision is made until the incremental population has been reviewed.
Negative / accepted costs
- The
reasoncontract moves to a new model.conversion_autoapproval_evaluationkeeps emitting exclusion-rule hits unchanged — it does not grow a fourth rule class, which was the cost of the rejected option C. What it does cost: the 33 unit tests assertingreasonverbatim now describe outputconversion_approval_decisionproduces, so they move with it, and they become a contract constraining future changes to the string's format. - Precomputed baselines are new models to schedule, monitor and back-fill — one per signal family that needs one. Cheap to run on one cadence, but not free: an incremental aggregate carrying daily moments has to be right about window boundaries, and getting that wrong is a silent numerical error rather than a failure.
- Stage 2 spends reviewer time on conversions that were already approved. Bounded and one-off per rule, but it is a real ask of the review team and needs their agreement, not their tolerance.
- Thresholds in code mean a tuning change is a deploy. Accepted while measuring, precisely because a silent mid-window change invalidates the measurement; revisited once a rule has a sole-reason precision measured at a shipped threshold and tuning is no longer part of establishing it.
Risks
- A rule validated on the wrong population. Stages 0 and 1 measure the overlap population only. Shipping on that evidence alone is the most likely way this goes wrong, which is why stage 2 is a decision rather than a recommendation.
reasonoverflow truncates silently.varchar(512), and thecastaroundLISTAGGcuts rather than raising — so the failure mode is a corrupted attribution string, not an error anyone sees. Bounded today by ~7 dimensions × 3 tiers and an observed max of 81 characters; a catalog of heuristics moves toward the ceiling. Needs a width check before the catalog grows, and it is worth doing before rather than after, since nothing will announce it.- Label return completeness, which matters more than review capacity. Every precision number in
this ADR depends on verdicts coming back for the conversions that were held. Measurement over
2026-08-06 to 08-22 found
report_datecarrying only 9 distinct values across 28 days, with a seven-day gap between 08-10 and 08-17. The gaps sit interior to the window rather than at the recent edge, so it is not review maturity. No cause is established here, but it points at the return path rather than reviewer throughput, and one question to the review team settles it: is the sheet returned once per report run, or in batches? If batches, the empty days are explained. If it is meant to be daily, verdicts are being lost and every precision figure in this document needs a day-mix caveat. - Review capacity is real but is not currently the binding constraint. The review team reports clearing a day's queue in about two hours, so a heuristic's cost is better argued as opportunity cost than as capacity — which is what the payout floor above is for.
- Two populations, two sets of numbers. The figures in this ADR do not all describe the same thing, and conflating them would be easy. The fraud report carries roughly 1,050–1,100 conversions daily with about 19% labelled fraud; it is a five-day window and includes rows the engine never held. The single-rule hold population in the actionable band is roughly 800 holds a day at 0.2097–0.2332 precision. Neither number is wrong and they are not in conflict; any measurement that quotes one should say which.
- Abstention that is invisible. Decision 6 makes the safe choice in three situations — thin data, a stale baseline, a failed model — but in the first two a heuristic that stops firing looks identical to one that found nothing. Same-cadence incremental makes staleness rare; insufficient support is the one that will happen routinely, on every new offer. Both need to be countable, not just safe.
- Loop latency. Anything added inline is paid ~144 times a day. A predicate that is cheap on a normal day may not be on a burst day.
NOTES
References
- Linear PEX-485 — this ADR
- Linear PEX-469 — auto-approval adjustments for heuristics in production
- ADR-0073: Config-as-Data Reference Tables — the rules substrate heuristics consume, and the fail-open lesson from Incident 203
- ADR-0072: Redshift Streaming Ingestion — rejected a real-time datastore for live fraud checks; the batch design here stays inside that
adaction/tune_fraud/autoapprove_conversions.py— the DAGdbt/dbt-adaction/models/conversion_review/autoapproval/conversion_autoapproval_evaluation.sql— the evaluation model, which keeps emitting exclusion-rule hits unchanged- PR #216: docs(ADR): heuristic rule layer for conversion auto-approval
Follow-Ups
-
ADR: heuristics vs ML — why both detection paths exist. Carries the argument decision 1 depends on, and the cross-path decisioning layer. Also has to reconcile
docs/diagrams/conversion-postback-flow.md, which currently shows the model emitting approve/reject/hold directly. -
ADR: technical architecture for Path B. Training, shadow serving, the risk-score contract, and what MLflow is actually used for on this model.
-
Converting each indicator from a model into a macro. Decision 4 sets the pattern and one indicator already follows it. Every remaining indicator needs its rule lifted out of its fraud-signal model into a macro before the pending path can call it. This is the bulk of the engineering between this ADR and a heuristic that actually runs, and it is per-indicator work, not a single migration.
-
Thresholds for every indicator. A whole piece of work rather than a detail of this one. Per indicator, a sweep across candidate thresholds over historical reviewed conversions, reported as trigger volume, sole-reason volume, precision, incremental fraud and payout exposure — then presented to the conversion review team as a capacity trade-off rather than a proposed number, and recorded as dbt vars once they choose an operating point.
-
Defining the actions above hold, with the conversion review team. Whether ranking and routing are distinct actions, what each means operationally, and what the review queue would need to carry them. The decisions are theirs, not ours.
-
Persisting rule-level attribution per run — a blocker on the velocity family, not a nice-to-have. Decision 2 already produces the (conversion, rule) long-form relation for heuristics; the
blocksCTE is the same shape for every other rule class. Persisting both per run gives per-rule trending and overlap without parsingreasonback apart.This is not merely retroactive drift in trigger counts. Measured: 0.00% of holds in the window were scored exactly once — the under-15-minute and 15-to-60-minute buckets returned no rows at all. The shortest hold was 71 minutes, the median 1,400, the longest 19,810 (~13.75 days). At a ten-minute cadence that is roughly 7 re-scoring cycles as a floor and about 140 at the median.
So latest-state attribution is not slightly lossy at the edges; re-scoring is the normal case. For any predicate whose value changes across a conversion's pending life — velocity most obviously, since a player's 10-minute count is a different number an hour later — a rule's contribution can vanish from the record entirely, and the trigger count for it is unreliable rather than approximate. The velocity family should not reach stage 3 until this lands. See the note in decision 5 on which predicates stage 3 can currently measure.
-
Threshold governance. Revisit code-versus-data once thresholds stop moving.