Authenticating...
Skip to main content

0077: Heuristics and ML as Parallel Detection Paths

STATUS​

Accepted

Records why Conversion Risk Intelligence maintains two detection paths rather than one, and how they converge. The technical architecture of each path is decided separately — see Follow-Ups.

CONTEXT​

The problem​

The Conversion Risk Intelligence architecture asserts two parallel detection paths on a shared fraud-intelligence foundation: behavioural heuristics, and an ML risk model. What it does not record is why both, what each is for, or how they coexist.

The absence has a cost, and it is already visible. Without a stated position, heuristics read as scaffolding to be removed once the model is ready. Every subsequent decision about the heuristic catalog then looks temporary, and the natural question on any heuristic ticket becomes "why are we building this if the model is coming?" That framing is wrong, and it is cheaper to correct here than to re-argue on each ticket.

What each path is​

A heuristic is a named, thresholded predicate over a conversion and its history — "the click country is outside the offer's targeting", "the goal completed faster than any plausible player". It encodes a hypothesis somebody had.

An ML risk model learns a ranking over the same features. It finds combinations nobody wrote down, and it produces a score rather than a reason.

Considered options​

  1. Heuristics as a stepping stone, retired once the model ships. The default assumption when both exist. Rejected below — the two fail differently and change at different speeds, and losing the heuristic path costs capabilities the model does not replace.

  2. Heuristics only. Cheapest and most explainable. Rejected: it can only ever catch mechanisms someone already articulated, which is a poor fit for an adversary that changes faster than the catalog does.

  3. ML only, deciding directly. What docs/diagrams/conversion-postback-flow.md currently depicts. Rejected in decision 6.

  4. Two paths converging on a policy layer. Chosen.

DECISION​

1. Both paths are permanent, and neither is a phase of the other​

Heuristics are not scaffolding and the model does not retire them. The case rests on four differences that do not go away as the model improves.

  1. They fail differently. A heuristic fails by being wrong about a rule you can read. A model fails by being wrong in a way nobody wrote down. Running both means a failure in one is visible against the other, rather than being the only account of what happened.

  2. They change at different speeds. A heuristic changes in hours through a reviewed diff. A model changes through retraining, revalidation and a deployment. When a fraud pattern appears on a Tuesday, the hours path is the one that responds.

  3. They explain differently. A reviewer can act on "held because the click country is outside the offer's targeting". Nobody can act on "0.62". The review team is the end user of both paths, and one of them produces something they can adjudicate.

    To be explicit, because the point above reads more strongly than intended: a score does not need a heuristic to fire alongside it before it can produce an action. A high score can hold a conversion on its own — that is precedence rule 4 below. What a bare score cannot do is be the explanation handed to the reviewer. The obligation that creates falls on the decisioning layer, which carries the score's attribution — the features that drove it — alongside the score, not on the heuristics to corroborate it. Requiring a rule hit before a score could act would throw away exactly the cases the model exists to find: the ones no rule describes.

  4. Heuristics are the baseline the model has to beat. Without an interpretable floor, "the model is good" has no denominator. That role does not expire.

2. The division of labour​

Heuristics cover known mechanisms — hypotheses someone has articulated and can defend. They gate directly, because a named rule hit is something a human can review.

The model covers residual risk — what remains after the known mechanisms, including combinations of weak signals no single rule would catch. It ranks rather than decides.

This is why heuristics stay deliberately independent of one another, one threshold per indicator with no combination logic: combining weak signals into a single verdict is the model's job, and a heuristic that tried would be an unexplainable model with none of the machinery.

3. One feature foundation serves both​

Both paths consume the same point-in-time-safe features from Redshift, dbt and Airflow. There is no second pipeline for ML.

This is the decision that makes the other three affordable. Two feature pipelines would mean two definitions of the same signal, drifting apart, with no way to tell whether a disagreement between the paths was a real disagreement or a plumbing difference.

4. Detection produces evidence; policy produces the action​

Neither path decides what happens to a conversion. A heuristic emits a named rule hit; the model emits a score with its model and feature versions. The decisioning layer turns evidence into an action.

Keeping this boundary is what allows a threshold to change without retraining, a model to be swapped without renegotiating business rules, and a score to be recorded while having no effect at all — which is what shadow mode is.

5. The decisioning layer combines evidence by ordered policy, not by a score​

The layer has to reconcile discrete, named rule hits with a continuous score. The obvious move is to normalise them into one weighted number. That would be a mistake here:

  • it makes every decision unexplainable to the reviewer being asked to adjudicate it;
  • thresholds stop being interpretable, since the same number means different things depending on which inputs produced it;
  • it couples the rules together, so tuning one changes the behaviour of all of them;
  • it destroys per-rule attribution, which is how we measure whether any individual rule earns its place.

Instead, evaluate ordered policy rules against the evidence and take the first that matches. Illustrative, not final:

PrecedenceConditionAction
1A hard exclusion rule hitHold
2Payout at or above the ceilingHold
3Any heuristic hit, above the payout floorHold
4Risk score in the high band, above the payout floorHold
5Risk score in the medium band, with any rule hit, above the payout floorHold
6Nothing matchedAuto-approve

Which rules carry the payout floor, and why it is not all of them. Rules 3, 4 and 5 do; rules 1 and 2 do not, and the split is principled rather than incidental.

The floor exists on an opportunity-cost argument: most traffic under the floor is never reviewed at all, so a rule firing below it does not add a variable to a population under review — it adds a population nobody looks at, displacing attention from the band where the money is. That argument is about what reviewer attention is worth spending on, and it is indifferent to which mechanism flagged the conversion. A fifty-cent conversion held by a high risk score costs the same attention, and returns the same negligible exposure, as one held by a heuristic. So the score path inherits the floor rather than escaping it.

Rules 1 and 2 are different in kind. Rule 2 is a payout rule. Rule 1's payout semantics already live in the exclusion tiers, where an always rule deliberately fires at any payout because somebody decided that offer or device warrants review regardless of value. That is a human judgment about a named thing, not a statistical one about a population, and the opportunity-cost argument does not override it.

So the line is between discretionary detection — rules we add because we believe they help, whose cost has to be argued — and deliberate exclusion, where the cost was already accepted when the rule was written.

The score floor is its own dbt var rather than a reference to the heuristic one, starting at the same value. Same queue and same opportunity cost today, but if the score proves more precise than a single heuristic hit it may justify reaching lower, and that should be a threshold change rather than a redesign.

Every action here is hold or auto-approve, because those are the two that exist. Richer actions are deliberately not named: whether the queue has a notion of priority at all, what it would mean operationally, and whether routing is separate from ranking are the review team's decisions, not ours to assume in a table. The Path A ADR takes the same position and tracks the same follow-up. What the ordering gives us regardless is attribution — rules 4 and 5 hold for different stated reasons than rules 1 to 3, so the record distinguishes a score-driven hold from a rule-driven one even while the action is the same.

Why rule 5 exists, and why it is a hypothesis rather than a finding. A medium score on its own is not strong enough to be worth review capacity; a medium score with independent corroborating evidence plausibly is. The two signals come from different mechanisms — one learned from many weak features, one a named condition somebody wrote down — so agreement between them is worth more than either alone. There is measured support for the general shape: in the current single-path data, conversions tripping two rules run far higher precision than those tripping one. There is no measurement yet for the score-plus-rule combination specifically, because no score is in production.

That makes rule 5 exactly what shadow mode is for. Scoring without acting produces the comparison — medium-band-with-a-rule-hit against medium-band-alone — before the row is granted any authority. If the corroboration effect is not there, the row does not ship.

Every decision names the policy rule that produced it. Bands tune independently of exclusion rules. A missing score means rules 4 and 5 simply do not match, and 1 to 3 still decide — absent evidence abstains rather than failing. Shadow mode becomes a configuration, not a code path.

The layer emits one record per conversion: the decision, the policy rule that produced it, the evidence present at the time, and the policy version. That record is what makes "why did this happen" answerable afterwards.

This is an evolution, not a new system. It is what conversion_autoapproval_evaluation and ConversionApprover become when they gain precedence ordering and the extra evidence columns.

6. This supersedes the ML-decides framing in the postback flow diagram​

docs/diagrams/conversion-postback-flow.md currently routes conversions from an ML model straight to three outcomes: high confidence to approved, flagged fraud to rejected, uncertain to held for review. That is the model deciding, and it contradicts decision 4.

Its staging is not superseded — Stage 1 in that document is "ML Integration (Shadow mode alongside Tune)", which agrees with the position here. What changes is outcome ownership: the model produces a score, the decisioning layer produces the outcome.

The Rejected (Discarded/Logged) outcome is also not admitted. Both technical ADRs exclude auto-reject, on the grounds that it sits beyond the boundary where a mistake can be undone by a reviewer.

How the paths converge​

CONSEQUENCES​

Positive​

  • Heuristic work stops being provisional, and the catalog can be invested in on its own terms.
  • A disagreement between the paths is informative rather than confusing, because both read the same features.
  • The review team gets an explanation they can act on for every held conversion, whichever path produced it.
  • Shadow mode for the model needs no separate code path — it is a policy configuration.

Negative / accepted costs​

  • Two detection paths are two things to maintain, monitor and explain. That is the cost of the capability, not an oversight.
  • The decisioning layer becomes a component with its own policy version, its own tests and its own review surface.
  • Ordered policy is less expressive than a learned combination. That is deliberate, and it is where the model earns its place rather than a limitation to remove.

Risks​

  • The ladder is proposed, not agreed. The action set above has not been reviewed with the conversion review team, who own the queue every action lands in. It is a starting point for that conversation, not a decision made on their behalf.

  • Policy sprawl. Ordered rules are readable at six and unreadable at sixty. The distinction that matters if the list starts growing is who authors the combination, not where it runs — both the policy and the risk model are dbt models in the warehouse, so "in the model rather than in policy" is not a statement about location. Policy is a short ordered list a person writes and reviews as a diff, where every row is explainable on its own. The risk model learns weights over many weak features from labelled data. A policy list growing toward sixty rows means someone is hand-fitting a classifier in SQL — badly, without a fitting procedure and without validation — and that combination should move into the learned model, with policy staying short and acting on its output.

  • Divergent feature definitions. Decision 3 holds only while both paths genuinely read the same models; a convenience copy on either side would quietly undo it.

  • The model is selected on one population and applied to another, and the paths are not equally exposed to it. Labels come from manually reviewed conversions, which is what the existing rules chose to hold — not a sample of production. Offline metrics therefore describe performance on the reviewed population, while the score would be applied to every pending conversion, and the gap between those is not measurable offline by construction: the conversions the model would newly flag have no verdicts, because nobody looked at them.

    This is asymmetric between the paths, which is part of why both exist. A heuristic's precision is measured on exactly the population it holds, and the incremental population it reaches is addressable by sampling it and asking for review — the sampled back-review stage in the Path A ADR. A model's offline metrics have no equivalent, which is the structural reason its first production milestone is shadow scoring rather than a decision: shadow is the instrument that measures the gap, and there is no offline substitute for it.

    Treated at length in the Path B ADR, which carries the shadow-first rollout and the conditions for leaving it. Recorded here because it bears on the relationship between the paths rather than only on the model.

NOTES​

References​

Follow-Ups​

  • ADR: technical architecture for Path A. How heuristics are computed and reach the approval decision. In review.
  • ADR: technical architecture for Path B. Training, shadow serving, the risk-score contract, and what MLflow is used for on this model.
  • Aligning the action ladder with the conversion review team. The decisions in it are theirs.
  • Updating conversion-postback-flow.md once this is accepted, so the diagram and the ADRs stop disagreeing.