Authenticating...
Skip to main content

0078: Technical Architecture for Path B (ML Risk Model)

STATUS​

Accepted

Records the target architecture for the ML detection path. Nothing in the DECISION section is built. Decision 1 carries an explicit gate that has not yet been run.

CONTEXT​

The problem​

The heuristic path has an ADR. The ML path does not, and the gap is not that nobody knows which tools to use — it is that "we will use MLflow" has been said often enough to sound settled while meaning something different to each person saying it.

That is the actual decision this ADR owes. Whether to track model evolution is not in dispute. What MLflow is for on this model, what it is not for, where SageMaker starts, and what PyCaret is allowed to become are all undecided, and an ADR that recorded "we adopt MLflow" would leave every one of them open.

Where the ML path is today​

Further along than it may look from outside, which is worth stating precisely, because the gap this ADR addresses is not a modelling gap.

What exists. A versioned training dataset with an executable feature contract. Rolling-origin temporal evaluation rather than random cross-validation. A baseline leaderboard across model families. A determinism harness that establishes per-model noise floors, so a margin can be read against measured variance rather than asserted. And a leaderboard interpretation — margin, SHAP attribution, per-month stability, operating points — that recommends a model family against a reading rule committed before the numbers were computed.

What does not exist. Experiment tracking, a model registry, lineage from a score back to the data that produced it, a production estimator, and any serving path at all.

So the asymmetry is the point: the modelling work has been careful, and none of its output is currently reproducible by anyone who did not run it. Runs live in notebook outputs and in the memory of whoever executed them. Two engineers cannot answer "which dataset produced this leaderboard" without asking each other. Every decision below exists to close that gap rather than to improve the model.

Decision drivers​

  1. A score has to be traceable to what produced it. Model version, feature version, training window. Without that the score cannot be audited, and a fraud decision that cannot be audited cannot be defended to a publisher.
  2. Experimentation must stay fast. The baseline work is not finished. A platform that makes iteration slower will be routed around.
  3. Operational burden has to match the team. The ML team is small. Infrastructure we run ourselves is infrastructure we maintain instead of doing modelling.
  4. The model must not acquire decision authority by default. Detection produces evidence; policy produces the action. That separation is the parent ADR's, and this one has to implement it rather than restate it.
  5. Batch is sufficient, and staying batch keeps an earlier decision intact. ADR-0072 rejected a real-time datastore for live fraud checks. A design needing sub-second reads reopens it.
  6. Nothing about the experimentation tool should become load-bearing. PyCaret earns its place today; it should be replaceable tomorrow without touching lineage or serving.

Considered options​

A: self-hosted MLflow. Best portability and a good developer experience, but the company owns infrastructure, upgrades, authentication, networking, persistence, availability and security. Rejected: that operational work returns little at our current team size, and driver 3 is the whole argument.

B: SageMaker-native, no MLflow. Strong production capability. Rejected as premature — adopting the ecosystem now means committing to Pipelines, automated retraining, a Feature Store and dedicated endpoints before training and serving requirements are stable, and it weakens driver 6 by making the platform the interface.

C: Databricks. Raised in review as the serious contender, and it is one: managed MLflow plus accelerators for exactly the experimentation patterns we would otherwise assemble. Rejected on fit rather than capability. The data foundation is Redshift, orchestration is Airflow, and the warehouse is where every feature already lives; adding a second compute platform to host one model is a larger commitment than the model currently justifies. Worth reopening if the ML footprint grows beyond this initiative.

D: managed MLflow on SageMaker AI. Chosen. AWS provides MLflow as a managed service, so the lifecycle interface is MLflow — portable, familiar, already integrated with PyCaret — while AWS owns the operational layer. SageMaker capability is then adopted piecewise when a requirement justifies it, rather than wholesale.

Comparison​

A: self-hostedB: SageMaker-nativeC: DatabricksD: managed MLflow
PyCaret integrationExcellentModerateExcellentExcellent
Platform maintenanceHighLowLowLow
Portability of lineageHighLowerModerateHigh
AWS / warehouse fitModerateExcellentModerateExcellent
Commitment incurred nowLowHighHighLow

DECISION​

1. Managed MLflow on SageMaker AI is the lifecycle interface, subject to a POC gate​

MLflow is the standard interface for experiments, artifacts, lineage and model versions, run as the AWS-managed service rather than infrastructure we operate.

This is a recommendation with a gate, and the gate has not been run. Stating it as settled would be dishonest. Before it becomes the platform standard, one existing Conversion Risk Intelligence experiment goes through the whole path — dataset, PyCaret run, managed MLflow, S3 artifacts, registered candidate model — and has to demonstrate:

  • the team can access and compare runs;
  • metrics, parameters and artifacts persist;
  • git commit, dataset version and feature-set version are traceable from a run;
  • fraud-specific metrics can be logged, not just the sklearn defaults;
  • a candidate model can be registered;
  • the experiment reproduces from the repository;
  • IAM and access patterns are workable;
  • cost and maintenance overhead are acceptable.

If the POC fails on operational overhead, option A is the fallback and this decision is revisited. Everything below assumes the gate closes.

2. The division of labour, stated concretely enough to build from​

The recurring confusion is treating these as alternatives. They occupy different layers.

OwnsExplicitly not for
PyCaretBaseline creation, model-family comparison, rapid iterationAnything in production. Not an architectural dependency
MLflowExperiments, runs, parameters, metrics, artifacts, dataset and feature metadata, model versions, the registryTraining compute. Serving. Orchestration
SageMakerManaged MLflow hosting, IAM, S3 artifacts, and later training and batch inferenceBeing the interface engineers work against
AirflowScheduling the feature build, the scoring run and the comparisonModel lifecycle state

The load-bearing consequence is that MLflow is the interface and SageMaker is the substrate. An engineer logs to MLflow; where it runs is an operational detail. That is what keeps driver 6 true: a future experiment in scikit-learn, XGBoost or PyTorch logs to the same place, and PyCaret can be dropped without touching lineage.

PyCaret's boundary is worth naming precisely, because it is easy to cross by accident. It answers what should we build — which family, which preprocessing, which features earn their place. It does not answer what should production run. The selected family is reimplemented deliberately as production estimator code. A PyCaret pipeline is not promoted to a registered model because it won an experiment.

3. Every run logs the same three categories, or the run is not evidence​

A leaderboard without provenance is an anecdote. Each experiment records:

Reproducibility — git commit, dataset version and training window, feature-set version, model parameters, random seed, decision threshold.

Offline performance — average precision as the primary metric given class imbalance, ROC-AUC, precision, recall, F1, confusion matrix, and the per-month breakdown rather than only the pooled figure. Average precision is the only precision-recall summary logged, deliberately: it is the step-wise estimator of the same curve a trapezoidal PR-AUC interpolates, and the interpolation reads optimistically under heavy imbalance. Logging both would put two numbers on one quantity and invite an argument about which to quote when a margin is close.

Business and risk performance — fraud leakage, false-positive and false-negative rates, approval rate, review volume, payout exposure.

Leakage is not a second name for the false-negative rate, and the distinction is the reason both appear. False positives and negatives are model-level: counts against reviewer verdicts, over the labelled reviewed sample, at a stated threshold. Fraud leakage is system-level: fraud that reached payout after deterministic rules, heuristics, the model and human review had all had their turn, denominated in exposure rather than in events, and over all traffic rather than the reviewed subset. A model can improve its false-negative rate while leakage is flat — if policy never acted on the scores, or if the fraud it now catches was already being caught by a rule. Neither number answers for the other.

The third category is not yet defined well enough to log, and that is recorded here rather than left implied. What is above is the intent, not a specification — the denominators, the measurement window and what "cost" means operationally all still need the review team's input. Until it exists, model selection is being made on offline metrics alone, which is a known limitation and not a neutral one: the best fraud model is not the one with the highest AUC. Selection optimises a trade-off between fraud prevented, good conversions affected, and manual review volume, and two of those three are currently unmeasured.

When they are defined, they are the heuristic path's definitions rather than new ones. Path A's measurement framework already has a first pass at two of the three terms. Fraud prevented is covered three ways: a measured baseline for what the marginal reviewed conversion is currently worth, a gate stated as a required integer count rather than a rate — because rounding crosses the boundary — and a dollar-of-confirmed-fraud-per-hold comparator, which is the one that actually maps to the leakage KR. Review volume is covered by a review-cost side derived from the reviewer queue. Good conversions affected is the term neither path has yet, which is why it is the half of the trade-off still missing rather than one more thing to define later.

A score threshold is the same decision as a rule with a continuous knob instead of a binary one: at what point does the marginal hold earn its review. If Path B invents its own vocabulary for that, the two paths stop being comparable to each other, and that comparison is the one the parent ADR has to be able to make.

This starts at the decision 1 POC — not before it, and not after the candidate is chosen. There is no tracking server today, so nothing is being logged and nothing is being lost: reproducibility currently rests on a committed feature contract, a fixed seed, a determinism harness and notebook outputs in the repository, which is exactly the manual arrangement this decision replaces. The POC runs against an existing experiment, so the first run logged is one that has already been run, and from that point every run of record logs rather than only the ones someone remembers to. Earlier runs are not backfilled — re-executing them under the harness is cheap and produces provenance that is true rather than reconstructed.

All three categories rest on labels that are currently incomplete, and that is a prerequisite rather than a caveat. Reviewer verdicts are being lost in ingest, and they are lost as whole reports rather than as a random sample of conversions. Measurement on the heuristic side puts review coverage near 0.55 over a sampled window (holds from 2026-08-06 to 08-22, measured before any ingest fix), and finds it behaves day-level rather than rate-like: some days sit at zero and others near one, tracking which reports ingested. The figure is a measurement of a period, not a constant, and it should move once the ingest gap closes.

Nothing here stops on that — the POC and model selection both proceed on the labels that exist — but two things follow. Every figure logged under offline and business performance is computed over a sample with whole days absent, so it is usable and not complete. And coverage near 0.55 inflates every measurement timeline by roughly 1.8x, so the windows both paths are planning around should shrink when the gap closes, not grow.

That second point is about the calendar time and hold volume needed to accumulate a given number of reviewed conversions, not about the target itself. The n a measurement needs is set by the precision it has to reach; improving coverage makes that n arrive sooner rather than making a smaller n acceptable. The alignment behind the whole reading is worth confirming directly before either path states it as established.

4. The model emits risk, never a decision​

The model produces a score and its supporting evidence. It never emits approve or reject. The decisioning layer in the parent ADR turns evidence into an action, and this is the record the layer consumes:

FieldMeaning
tune_event_idThe conversion scored
prediction_timestampWhen the score was produced
model_name, model_versionThe registered model, resolvable in MLflow
feature_versionThe feature contract the score was computed under
training_dataset_versionLineage back to the data that produced the model
risk_scoreThe continuous output
risk_bandThe score bucketed, if bands are defined; policy consumes the band, not the raw score
attributionThe named signals that drove this score, from a bounded vocabulary — what makes it explainable to a reviewer, and countable
execution_metadataRun identifier, scoring job, duration

attribution is a requirement, not an enrichment. The parent ADR's explainability argument is that nobody can act on "0.62". The obligation that creates lands here: a score arrives with the features that produced it, or the decisioning layer has nothing to put in front of a reviewer.

It also has to be enumerable, and that is a measurement requirement rather than a presentation choice. Free text or an unbounded top-k feature list makes a score explainable to a human and useless to an analyst: "scores where this signal was the only reason" stops being a query. The heuristic path gets that question for free, because its reason is an alphabetically ordered list of rule names and sole-reason is an equality test. Path B needs the equivalent — a bounded set of named signals from a versioned vocabulary, ordered deterministically — or there is no per-reason precision for the score, and no way to compare a score band against a rule.

There is deliberately no recommended_action column. Adding one would put policy inside the model's output and undo the separation this decision exists to protect. If a policy verdict needs storing, it belongs on the decisioning layer's record alongside the policy version that produced it.

5. Scoring is batch, which keeps ADR-0072 intact​

Conversions are scored in batch on the same cadence the decisioning layer runs. Features are read from Redshift, which is already the offline feature source and already correct at that grain.

No feature store, no online store, no real-time endpoint. ADR-0072 rejected a real-time datastore for live fraud checks and a batch design stays inside that ruling. A feature store becomes arguable when there is a concrete need — synchronous inference, multiple online consumers, strict freshness — and none of those exist. Introducing one now would be platform complexity ahead of a requirement.

6. Shadow first, and the conditions for leaving shadow are stated in advance​

The first production milestone is scoring that changes nothing. The model runs, writes risk records, and the decisioning layer ignores them.

Shadow is not a formality. It is the only way to observe score distribution, calibration, coverage, agreement and disagreement with heuristics and with reviewer outcomes, feature correctness in production rather than in a notebook, and runtime reliability — before any of it can cost a publisher a conversion.

What shadow will not show, and the instrument that fixes it. Scoring the whole pending population gives a score distribution over far more traffic than the heuristic path ever sees. It does not give labels on the conversions the engine auto-approves, because those are never reviewed — so the population where a missed fraud is most expensive is exactly the population with no ground truth. That is the same gap Path A's staged rollout already exists to close, and the sampled back-review mechanism should be scoped once and shared rather than discovered twice. A score-banded version is the same instrument with a different selection rule, and sizing it is work Path A has already done.

Before the score influences a decision:

  • the POC gate in decision 1 has closed;
  • shadow has run long enough to cover a full review cycle, with the score compared against actual reviewer verdicts rather than against the offline holdout;
  • production feature values have been reconciled against the training-time definitions, because a silent divergence there is the failure mode that offline metrics cannot see;
  • an operating point has been chosen from the PR curve and translated into operational consequences — fraud caught, legitimate conversions affected, review volume added;
  • the business and risk metrics in decision 3 exist, so the trade-off is measured rather than asserted;
  • sampled back-review is running over auto-approved traffic, so the missed-fraud side of the score is measured rather than assumed;
  • the review team has agreed to the review volume the operating point implies.

The holdout is scored once, after the selection procedure is frozen. Scoring it repeatedly to choose a family, preprocessing, features or thresholds turns it into another validation set, and there is then no unbiased estimate left.

7. What is deliberately not being built yet​

Naming these matters, because each is a reasonable thing to want and none is justified now: SageMaker Pipelines, automated retraining, a Feature Store, real-time endpoints, and automated model monitoring beyond the shadow comparison.

Each becomes arguable when a requirement demands it. Building them ahead of that is the specific failure this ADR is shaped to avoid — an ML platform that is impressive and unused while the model it exists to serve is still being selected.

CONSEQUENCES​

Positive​

  • A score becomes traceable to the model, features and data that produced it, which is what makes it auditable and therefore defensible.
  • The lifecycle interface is portable. If the managed service stops fitting, MLflow moves.
  • PyCaret stops being a dependency and becomes a tool, so a model-family change does not become a platform change.
  • Model and policy stay separable, so thresholds move without retraining and policy can be experimented with independently.
  • Shadow mode makes the first production milestone one that cannot cost a publisher anything.

Negative / accepted costs​

  • Managed MLflow is a running cost for a model that is not yet in production, incurred before the return is demonstrated.
  • Reimplementing the selected family as production code is real work that a PyCaret artifact would appear to save. It is the price of not deploying an opaque pipeline.
  • Batch scoring means a conversion is scored on a cadence rather than on arrival, so risk evidence is available later than a rule hit is.
  • Logging discipline is a convention until something enforces it. A run missing its feature version is not evidence, and nothing currently rejects it.

Risks​

  • The POC gate is skipped under time pressure. The likeliest failure, because the recommendation reads as settled. Decision 1 states the gate; treating it as done because the ADR merged is the mistake to guard against.
  • Training and serving diverge. Features computed one way for training and another way for scoring produce a model that looks good offline and behaves differently in production. The macro work in the Path A ADR is the mitigation, since a served model becomes a third caller of one definition rather than a reimplementation.
  • Business metrics stay undefined. Decision 3 depends on them and they do not exist. If they remain missing, model selection continues on offline metrics and the operating point is chosen without knowing its cost.
  • Selection bias in the labels is carried into production expectations. The model learns from manually reviewed traffic, which is a selected population. Offline performance does not translate directly to performance over all traffic, and shadow mode is the first honest measurement of that gap.
  • Label completeness gates every number in this document. Reviewer verdicts are currently being lost in ingest, as whole reports rather than as a random sample, so the loss is not random with respect to anything a day-level effect correlates with. Decision 3 names it as a prerequisite rather than a caveat. Treating the current figures as complete, rather than as usable pending the backfill, is the specific mistake to guard against.

NOTES​

References​

  • Linear PEX-487 — this ADR
  • Linear PEX-427 — the ML platform discovery this decision rests on
  • Linear PEX-600 — the review follow-up that added the enumerable-attribution requirement, the shared metric vocabulary, the shadow back-review dependency and the label-completeness gate
  • Linear PEX-567 — the ingest defect behind the label-completeness prerequisite in decision 3
  • ADR: heuristics and ML as parallel detection paths — the decisioning layer this model feeds, and the detection-versus-decision separation decision 4 implements
  • ADR: heuristic rule layer for conversion auto-approval — the macro pattern decision 4's training/serving parity depends on
  • ADR-0072: Redshift Streaming Ingestion — rejected a real-time datastore for live fraud checks; decision 5 stays inside it
  • PR #223: docs(ADR): technical architecture for Path B (ML risk model)
  • PR #224: docs(ADR): carry the Path B review feedback into 0078

Sequencing: what happens between here and a score that matters​

The decisions above describe a target. This is the path to it from where the work actually stands, and the ordering matters because two of these run in parallel and the rest do not.

Model selection is nearly done and is not blocked by anything here. The family recommendation is a gradient-boosted ensemble, with the choice between the top three not yet settled and the supporting evidence in review. Characterising the chosen candidate — segment behaviour, stability, operating points — follows it. Neither step needs MLflow, a registry or a serving path, and neither should wait for them.

The platform POC can start now, and should. Decision 1's gate runs against an existing experiment. It does not need the final candidate, and running it early is what stops the platform question from becoming the thing that blocks serving later. This is the parallel track.

Then, in order, and each genuinely depending on the one before it:

  1. Reimplement the selected family as production estimator code. Decision 2's boundary made concrete. This is the first step that cannot begin before selection concludes.
  2. Register it, with the lineage decision 3 requires — git commit, dataset version, feature version, seed, threshold. The first artifact that is reproducible by someone who did not train it.
  3. Build the scoring job and the risk-score record of decision 4, writing to a table nothing reads yet.
  4. Reconcile production feature values against training-time definitions. Deliberately before shadow rather than during it, because a silent divergence here is invisible to every offline metric and would otherwise be discovered by a wrong score.
  5. Shadow, under decision 6, with the exit conditions already written down.

Two things on the critical path are not modelling work and are not owned here. The business and risk metrics of decision 3 need the review team, and they gate step 5's exit rather than step 5's start. Label completeness gates the trustworthiness of every metric in this document, including the ones already measured. Both are tracked separately; neither resolves by building anything in this ADR.

Where infrastructure decisions live​

This ADR decides the shape of the ML platform — which layer owns what, what the contracts are, and what is deliberately not built. It does not decide how any of it is provisioned. That is intentional, but it should not leave a reader wondering whether infrastructure was forgotten, so:

Inherited, and not re-decided here.

Networking, corrected twice — and the second correction withdraws the first. The original revision of this section said the ML platform sits inside the shared VPC of ADR-0062. A later correction replaced that with a SageMaker account carrying its own VPC, interface endpoints, NAT gateways and a live peering connection to the data-warehouse account, provisioned by the aws-sagemaker-cdk-app repository.

That second version is the one that is wrong, and it is wrong about an estate that is not there. The data-warehouse VPC carries five active peering connections, every one of them created by Prod-DataWarehouseStack, and none terminates in a SageMaker account. The description was written from the archived CDK definition rather than from the accounts, and a definition is not a deployment.

The original claim was closer to true and is restored here with its mechanism named. The dedicated machine-learning-production account — created since, and the subject of the ML platform and deployment lifecycle ADR — has two VPCs shared into it through RAM, both of which hold active peering to the data-warehouse account. Network access is therefore inherited under ADR-0062 rather than built, which is what "networking is inherited, not a question here" was relying on. What is not inherited is a private path belonging to this workload, and nothing in this ADR needs one.

This is recorded rather than quietly fixed for the same reason the first correction was: the claim is load-bearing for anyone planning work against the platform, and two people have now planned against a version of it that was false.

Provisioning, and it is open in a sharper way than first recorded. No accepted ADR covers how AWS resources for a workload like this get provisioned. ADR-0035 adopts CDK but is Request for Comments and scoped to Identity Center rather than workload infrastructure.

The SageMaker estate itself was provisioned in CDK, by a repository that is now archived and read-only, and whose last published release predates this work by more than a year. Deployment from it fired on a published release, so nothing has been deployed from that definition in that time. Nor is the estate reachable: the account it targeted is absent from the organisation's account inventory, and no peering from it survives on the warehouse side. What the archived definition is good for is showing what a SageMaker estate needed when it was expected to host development work.

That makes the provisioning question narrower and harder at once. It is not "CDK or console" — the estate is already described in CDK. It is where that description now lives, whether the estate it describes is still deployed as written, and what replaced the repository when it was archived. Those belong to the POC in decision 1, alongside which account the tracking server lives in and who can reach it, whether its artifact store is exempted from the RemovalPolicy.DESTROY the existing SageMaker buckets carry, and whether the account's cost budget covers a tracking server at all.

Deferred until there is a model to serve. Three things will need their own decisions and none can be made usefully yet, because all three depend on a selected candidate that does not exist:

  • the promotion path from a registered candidate to production scoring, and what gates it;
  • what triggers retraining, which decision 7 defers explicitly;
  • what model monitoring watches, at what threshold, and who owns the alarm.

Writing any of these now would be the same mistake this ADR avoided by waiting for the platform discovery: a decision recorded before the evidence that would inform it.

Follow-Ups​

  • Run the POC and record the result against the eight criteria in decision 1. Until this happens the platform decision is a recommendation.
  • Define the business and risk metrics with the conversion review team. Decision 3 is incomplete without them, and they are the half of model selection currently missing.
  • Specify the risk-score table. Decision 4 defines the record; where it is written, at what grain, and with what retention is an implementation question this ADR does not settle.
  • Reconcile production and training feature values. The concrete test behind the training/serving risk, and it should run before shadow rather than after.
  • Decide what triggers retraining. Deliberately excluded from decision 7 for now, but a model in shadow will eventually need an answer that is not "when somebody notices".
  • Confirm the report-ingest alignment behind the coverage figure, jointly with Path A. Decision 3 states it as measured and not yet established, and neither path should harden it until the dates are checked directly.
  • Define the attribution vocabulary. Decision 4 requires it to be enumerable. Which signals are in it, and how it is versioned against the feature contract, is not settled here.
  • Scope the sampled back-review mechanism once, across both paths. Decision 6 depends on it and Path A has already sized it. Building it twice is the outcome to avoid.
  • Write the deployment and platform ADR this one defers to. Decision 7 names Pipelines, automated retraining, a Feature Store, real-time endpoints and model monitoring as not-yet, and decision 1 defers the end-to-end path to a POC gate. Review made the fair point that "not yet" is still a position worth recording in advance, and that those deferrals are where the team is least aligned. Naming them as out of scope here was a scoping decision, not an argument that they need no ADR.