Skip to main content

Command Palette

Search for a command to run...

I Finally Understand Data Drift: Your ML Model Didn't Break, The Data Changed.

Updated
•19 min read•View as Markdown

I used to think that once my model was trained on good data, the data problem was basically over.

Then I learned about data drift.

And suddenly I had a slightly uncomfortable question:

What happens when the data my model sees tomorrow doesn't look like the data it learned from yesterday?

The model hasn't changed.

The code hasn't changed.

Nobody broke anything.

But the data changed.

That's data drift.

And once I understood that, a lot of the monitoring and MLOps conversations I'd been seeing started making much more sense.

So… what exactly is data drift?

Let's start without the fancy terminology. Imagine I train a machine learning model using this kind of data:

Training data

Transaction amount:
$10
$15
$25
$32
$41
$50
...

The model learns patterns from this data.

Then I deploy it.

A few months later, the incoming transactions look more like:

Production data

Transaction amount:
$80
$120
$150
$210
$340
$500
...

The model is now seeing a different distribution of values.

Nothing necessarily went wrong with the model itself.

The data it receives simply changed.

That's the basic idea behind data drift:

Data drift happens when the distribution of incoming data changes relative to a reference distribution, such as the data used during training or data observed during an earlier period.

In production ML, we often compare a baseline distribution with a more recent production distribution to look for these changes. Google Cloud's model-monitoring documentation describes drift in terms of production feature distributions changing over time, while tools such as Evidently provide statistical tests and distance measures for comparing distributions.

And that sounds complicated until you visualize it.

Think about two distributions

Imagine we have one feature:

transaction_amount

During training, most transactions fall somewhere around $20–$100.

If we plot that distribution, we might get something like:

Training data

Frequency
  │
  │          ███
  │        ███████
  │      ███████████
  │    ███████████████
  │  ███████████████████
  └────────────────────────
      $10  $30  $50  $70  $90

Now our production data starts looking like this:

Production data

Frequency
  │
  │                         ███
  │                       ███████
  │                    ███████████
  │                 ███████████████
  │              ███████████████████
  └──────────────────────────────────
      $10  $50  $100  $150  $200

The shape has moved.

The distribution has changed.

That's the thing we're trying to detect.

The training distribution vs the production distribution

This distinction is probably the easiest mental model I've found.

Training distribution

This is the data your model learned from.

You can think of it as:

"What did the world look like when I taught the model?"

Production distribution

This is the data the model is seeing after deployment.

You can think of it as:

"What does the world look like now?"

If those two worlds start looking very different, you may have drift.

And this is one reason production ML is different from simply training a model and checking its test accuracy.

Your test set is a snapshot.

Production is a moving target.

A fraud detection example

Let's make this concrete.

Suppose we're building a credit-card fraud detection model.

Our training data contains features such as:

transaction_amount
transaction_time
merchant_category
location
device_information

The model learns patterns associated with fraudulent and legitimate transactions.

Maybe during training, fraudulent transactions commonly had certain combinations of:

  • transaction amounts

  • transaction times

  • locations

  • merchant categories

We evaluate the model.

The results look good.

So we deploy it.

But then something changes.

Maybe customers start using a new payment platform.

Maybe online shopping increases.

Maybe people start making larger transactions.

Maybe a new type of fraud becomes common.

Maybe fraudsters change their behavior specifically because they have figured out what older fraud detection systems look for.

Now the production data may no longer resemble the training data.

The model is still using the same learned parameters.

But the environment around it has changed.

That's where data drift becomes important.

But here's something I initially misunderstood

When I first started thinking about drift, I had this mental shortcut:

Drift detected = model broken.

That's not quite right.

A drift alert is not a magical message from the universe saying:

“Congratulations. Your model is useless now.” 😭

It means:

“Something about the data distribution changed enough that you should investigate.”

That distinction matters.

A model can experience some amount of distribution change and continue performing perfectly well.

For example, imagine the average transaction amount increases slightly because people are spending more during a holiday season.

That's a real change in the data.

But it doesn't automatically mean your fraud detector has stopped working.

The important question becomes:

Did the distribution change affect the model's ability to make useful predictions?

That's a much better question than simply asking:

“Did drift happen?”

Google's model-monitoring guidance similarly treats drift as something to detect and investigate; after identifying changes in feature distributions, engineers may need to examine the cause, adjust the data pipeline, or consider retraining.

Why does data drift happen?

There isn't one single cause.

Real-world data changes for all sorts of reasons.

1. User behavior changes

People don't behave exactly the same way forever.

For example:

Before:

Most transactions:
$20–$80

Later:

Most transactions:
$50–$200

The distribution changed because user behavior changed.

2. The environment changes

Sometimes the world around the system changes.

A business might enter a new market.

A new payment method might become popular.

A new product might launch.

A new regulation might change customer behavior.

A pandemic, economic event, holiday season, or other major event can also change patterns.

The model doesn't get a memo about any of this.

It just receives the new data.

3. The data collection process changes

This one is particularly sneaky.

Suppose your application originally collected:

age
location
device_type
transaction_amount

Then someone changes the application.

Now device_type is collected differently.

Or a feature starts returning different values.

Or a new category gets introduced.

Or missing values suddenly increase.

Your model might still be perfectly fine.

The pipeline feeding it isn't.

This is why data quality and data drift are closely connected, but they aren't exactly the same thing.

4. New populations enter the system

Imagine you train a model on customers from one region.

Later, your product expands into another region.

The new customers may have different:

  • spending patterns

  • demographics

  • behaviors

  • device usage

  • transaction sizes

The incoming distribution can change because your population changed.

5. Fraudsters adapt

Fraud detection is a particularly interesting example because the people you're trying to detect are not passive.

If attackers change their behavior, the patterns in your data can change too.

That creates a particularly interesting production problem:

Your model learned yesterday's patterns.

The attackers are trying to create tomorrow's patterns.

So even a model that performed well during evaluation can become less useful as the underlying behavior changes.

This broader problem is related to distribution shift and out-of-distribution generalization, which has been studied extensively in machine learning research.

Data drift vs covariate shift

This is where terminology starts getting annoying.

Because you'll often see terms like:

  • data drift

  • distribution shift

  • covariate shift

  • concept drift

  • dataset shift

and it can feel like everyone is inventing new names for the same thing.

They're related, but they aren't identical.

Let's break them down.

Distribution shift

At the broadest level, distribution shift means the data distribution relevant to your model changes between training and deployment.

You can think of:

Training world
      ↓
    Model
      ↓
Production world

If the statistical properties of those worlds differ, you have some form of distribution shift.

Research on domain adaptation has long studied what happens when training and target data come from different distributions.

Covariate shift

Covariate shift is a more specific situation.

The distribution of the input variables changes:

P(X)

but the relationship between the inputs and the target is assumed to remain the same:

P(Y | X)

So, roughly:

P_training(X) ≠ P_production(X)

but

P_training(Y | X) ≈ P_production(Y | X)

You don't need to memorize that equation immediately.

The intuition is more important:

The kinds of inputs you're seeing changed, but the underlying relationship between the inputs and the outcome hasn't fundamentally changed.

For example, suppose your fraud model normally sees transaction amounts between $10 and $100.

Later, most transactions are between $50 and $200.

That's a change in the input distribution.

If the relationship between transaction patterns and fraud hasn't changed, that is closer to the covariate-shift setting.

The distinction matters because different types of distribution changes can require different responses.

What about concept drift?

Concept drift is different.

Here, the relationship between the inputs and the target changes.

In simple terms:

The same kind of data no longer means the same thing.

Imagine that a certain transaction pattern used to be strongly associated with fraud.

Over time, legitimate customers start producing that same pattern.

Now the relationship between the feature and the target has changed.

That's much more concerning for the model.

You can visualize the difference like this:

DATA DRIFT

Input distribution changes

Training:
X → X → X → X

Production:
      X → X → X → X

versus:

CONCEPT DRIFT

Relationship changes

Before:
X ─────→ Fraud

Later:
X ─────→ Legitimate

The terminology gets more nuanced in the academic literature, but this mental model is useful when you're starting out.

So how do we actually detect data drift?

Now we get to the engineering part.

We need two things:

A reference

Something representing the distribution we're comparing against.

For example:

Training data

or:

Production data from last month

A current dataset

The newer data we want to inspect.

For example:

Production data from this week

Then we compare their distributions.

Conceptually:

Reference data
      │
      │
      ▼
Distribution A
      │
      │ compare
      ▼
Distribution B
      ▲
      │
Current production data

Google Cloud's current monitoring documentation describes this same general approach: establish a baseline distribution and compare newer production data against it, using statistical distance measures and alert thresholds.

Method 1: Visual inspection

Honestly, don't underestimate this.

Sometimes the first thing you should do is simply plot the distributions.

For a numerical feature, you might use:

  • Histogram

  • KDE plot

  • Box plot

  • Density plot

For example:

import matplotlib.pyplot as plt

plt.hist(reference_data["amount"], alpha=0.5, label="Reference")
plt.hist(current_data["amount"], alpha=0.5, label="Current")

plt.xlabel("Transaction Amount")
plt.ylabel("Frequency")
plt.legend()
plt.show()

You're looking for obvious changes.

Maybe the current distribution shifted to the right.

Maybe it became wider.

Maybe there are new outliers.

Maybe an entirely new cluster appeared.

This isn't necessarily enough for automated monitoring, but it's a fantastic way to understand what's happening.

And honestly, visualizing the data can make the whole concept click much faster than reading three pages of mathematical definitions.

Method 2: Statistical tests

For automated monitoring, we usually need something more systematic.

This is where statistical tests and distance measures come in.

Different tools support different approaches.

Evidently, for example, supports methods including:

  • Kolmogorov-Smirnov test

  • Wasserstein distance

  • Jensen-Shannon distance

  • Population Stability Index (PSI)

  • Kullback-Leibler divergence

  • Anderson-Darling test

  • Chi-squared and proportion-based tests for some categorical/binary cases

The appropriate method depends on the data type, sample size, and monitoring setup.

Let's look at a few of them.

Kolmogorov-Smirnov test

The KS test compares two distributions.

For example:

Reference distribution
        vs
Current distribution

It looks at the difference between their cumulative distributions.

For numerical features, it can help answer:

“Are these two samples plausibly coming from the same distribution?”

Evidently uses the two-sample KS test among its supported drift-detection methods.

One important thing to remember:

A statistical test result is evidence about distributional change.

It isn't automatically evidence that your model's business performance has degraded.

Wasserstein distance

This one has a more intuitive interpretation.

Imagine your reference distribution as a pile of sand.

Your current distribution is another pile.

The Wasserstein distance measures, roughly speaking, how much “work” it would take to move one distribution into the other.

You don't need to calculate it by hand.

The important idea is:

Small distance
    ↓
Distributions are relatively similar

Large distance
    ↓
Distributions are more different

Wasserstein distance is one of the statistical measures supported for numerical data in Evidently's drift tooling.

Population Stability Index

You'll also see PSI, especially in areas such as finance and risk modeling.

PSI compares how observations are distributed across predefined bins between a reference dataset and a current dataset.

Conceptually:

Reference:

0–20      ██████████
21–40     █████████████
41–60     ███████
61–80     ███

Current:

0–20      ████
21–40     ██████
41–60     ███████████
61–80     ██████████

The distributions have shifted.

PSI gives you a numerical way to quantify that shift.

Again, though:

Don't blindly memorize a threshold and declare:

“PSI > X means retrain immediately.”

Thresholds depend on your monitoring setup, domain, data and risk tolerance.

Evidently allows PSI and other statistical methods to be configured with thresholds rather than treating one universal number as the answer.

Here's where I think beginners can get confused

You detect drift.

You get:

🚨 DRIFT DETECTED

And immediately think:

“RETRAIN THE MODEL!!!”

Not necessarily.

Pause.

Investigate first.

Because there are many reasons why your data might have changed.

Maybe:

  • It's a temporary seasonal effect.

  • A new customer group joined.

  • A feature pipeline changed.

  • The data contains a bug.

  • A new category was introduced.

  • A real-world event changed behavior.

  • The model is genuinely becoming stale.

Those situations require different responses.

So I think a better production mindset is:

Drift detected
      ↓
Investigate
      ↓
What changed?
      ↓
Why did it change?
      ↓
Does it matter?
      ↓
Is model performance affected?
      ↓
Choose response

Not:

Drift detected
      ↓
RETRAIN!!!

😭

What should you do after detecting drift?

There isn't one universal answer.

The next step depends on what you discovered.

Case 1: It's harmless

Suppose transaction amounts changed because it's a holiday season.

The model's performance remains strong.

You might simply continue monitoring.

No emergency retraining required.

Case 2: It's a data pipeline problem

Suppose you discover that a feature suddenly contains mostly null values.

That's not necessarily a model problem.

It's a data problem.

You may need to fix the pipeline before touching the model.

Case 3: The model performance has degraded

Now things get more interesting.

If you have labels available, you can evaluate the model on recent production data.

Maybe:

Before:

Precision: 94%
Recall:    89%

Now:

Precision: 83%
Recall:    71%

Now you have stronger evidence that something is wrong.

You can investigate the cause and consider retraining or other model changes.

Google's ML guidance recommends continuous monitoring because serving-data changes can contribute to degradation, and suggests investigating drift and skew before deciding whether to adjust training data or retrain.

What if we don't have labels?

This is a really important practical problem.

Imagine your fraud detection system makes predictions today.

But you don't immediately know whether every transaction was actually fraudulent.

The true labels might only become available later.

So you can't simply calculate today's accuracy.

Instead, you can monitor things you can observe immediately:

Input distributions
Prediction distributions
Data quality
Missing values
Feature drift
Latency
Error rates

Then, when labels eventually become available, you can evaluate actual model performance.

This creates an important distinction:

Drift monitoring
       ≠
Performance monitoring

Drift can be an early warning signal.

It isn't the same thing as measuring model accuracy.

This is why monitoring needs context

Imagine you monitor ten features.

Nine look exactly the same.

One feature has significant drift.

Should you retrain?

Not necessarily.

You need to know:

  • Which feature changed?

  • How much did it change?

  • Why did it change?

  • How important is that feature?

  • Is the change temporary?

  • Is model performance affected?

  • Is the data pipeline working correctly?

This is where production ML starts feeling less like:

train()
predict()

and more like engineering.

You aren't just asking whether the model works.

You're trying to understand the system around it.

A simple drift-monitoring architecture

Here's the mental architecture I'm taking away:

                 ┌──────────────────┐
                 │  Reference Data  │
                 │   Training /     │
                 │  Previous Period │
                 └────────┬─────────┘
                          │
                          │
                          ▼
                   ┌─────────────┐
                   │ Distribution│
                   │   Baseline  │
                   └──────┬──────┘
                          │
                          │ compare
                          │
Production Data ──────────┤
                          ▼
                  ┌──────────────┐
                  │ Drift Test / │
                  │   Distance   │
                  └──────┬───────┘
                         │
                ┌────────┴────────┐
                │                 │
             No drift          Drift
                │                 │
                ▼                 ▼
             Continue        Investigate
             monitoring           │
                                  ▼
                          Check data quality
                                  │
                                  ▼
                         Check model performance
                                  │
                         ┌────────┴────────┐
                         │                 │
                      Healthy          Degraded
                         │                 │
                         ▼                 ▼
                     Continue          Retrain /
                     monitoring        improve

This is a much better mental model than:

“Drift means retraining.”

A small Python example

Let's make this concrete with a very simplified example.

Suppose we have a reference dataset and a current production dataset.

import pandas as pd
from scipy.stats import ks_2samp

reference = pd.read_csv("reference_data.csv")
current = pd.read_csv("current_data.csv")

feature = "transaction_amount"

statistic, p_value = ks_2samp(
    reference[feature],
    current[feature]
)

print(f"KS statistic: {statistic:.4f}")
print(f"p-value: {p_value:.4f}")

We're comparing the distribution of transaction_amount in the reference data with the current data.

You might see something like:

KS statistic: 0.1834
p-value: 0.0021

A small p-value can provide evidence against the assumption that the two samples come from the same distribution.

But here's the part I really want beginners to remember:

Don't turn a statistical test into a magic “model is broken” button.

The result tells you about the distributions.

You still need engineering and domain context to decide what the change means.

And in a real monitoring system, you'd also need to think about sample sizes, multiple features, alert thresholds, missing values, data quality, monitoring frequency and what action should follow an alert.

Why data drift belongs in MLOps

This is where Day 1 connects directly to Day 3.

In my first article, I talked about how MLOps isn't just:

Train model
↓
Deploy model
↓
Done

Data drift is one of the reasons why.

Your model is deployed.

People are using the system.

New data is constantly arriving.

And the distribution of that data can change.

So the production lifecycle becomes something more like:

Train
  ↓
Deploy
  ↓
Collect production data
  ↓
Monitor
  ↓
Detect changes
  ↓
Investigate
  ↓
Evaluate
  ↓
Improve
  ↓
Deploy again

The model becomes part of a loop.

Not a finished artifact.

The biggest mental shift for me

I think the simplest way I can explain data drift now is this:

Imagine teaching someone how to recognize cats.

You show them thousands of pictures.

Most of the cats are:

Indoor
Good lighting
Front-facing
High-resolution
Clear backgrounds

Then you send them into the real world.

Suddenly they're seeing:

Dark photos
Outdoor cats
Different camera phones
Unusual angles
Multiple cats
Partial images
Different environments

The person hasn't forgotten what a cat is.

But the world they're seeing is different from the examples they learned from.

That's the intuition behind distribution shift.

Machine learning models have the same basic problem, except they don't get to complain about it.

What I understand now

Before learning about data drift, I mostly thought about ML like this:

Good data
   ↓
Train model
   ↓
Good accuracy
   ↓
Deploy

Now I'm thinking:

Training data
      ↓
    Model
      ↓
Production data
      ↓
Is the data still similar?
      ↓
      NO
      ↓
Investigate
      ↓
Is performance affected?
      ↓
Maybe retrain

And that little question—

“Is the data still similar?”

—is a surprisingly big part of production ML.

Because your model doesn't know that the world changed.

It only knows what it learned.

One thing I don't want to do anymore

I don't want to see a drift dashboard showing:

🚨 DRIFT DETECTED

and immediately panic.

I want to ask: What changed?

Then: Why did it change?

Then: Does it matter?

And finally:

What should we do about it?

That feels like a much more mature way of thinking about ML systems.

Because the goal of monitoring isn't to produce scary red dashboards.

The goal is to help us understand what's happening to a system we care about.

Try this on one of your own ML projects

If you've already built an ML model, pick one numerical feature.

For example:

age
price
transaction_amount
temperature
income

Take your original dataset and pretend it represents your reference distribution.

Then create or collect a second dataset representing newer production data.

Plot the two distributions.

Ask:

  1. Do they look different?

  2. Which feature changed the most?

  3. Is the change actually meaningful?

  4. What could have caused it?

  5. Would the change necessarily hurt model performance?

  6. What would you monitor in production?

  7. What would make you investigate further?

  8. What would make you retrain?

That's the part of data drift I finally understand: Drift detection isn't the end of the investigation. It's the beginning.

And honestly, that's what makes production ML so interesting.

You aren't just building a model.

You're building something that has to keep making sense while the world around it changes.

Day 1 was about understanding MLOps.

Day 2 was about realizing that accuracy isn't the finish line.

Now Day 3 is about one of the reasons that finish line keeps moving.

The data changes.

And your model has to be ready for that.

I Finally Understand…

Part 3 of 4

A beginner-friendly, 10-day guide demystifying MLOps and production machine learning. From data drift and pipeline management to post-deployment model monitoring, learn how to bridge the gap between building a model and keeping it running successfully in the real world.

Up next

Model Drift vs Data Drift: What's Actually the Difference?

Your model is making predictions, the API is responding, nothing appears to be broken. But something has changed. The transactions reaching your fraud detection model look different from the ones it w