Why ML systems are fundamentally different from regular software — the hidden complexities of data drift, technical debt, and the three axes of change.
In traditional software, behavior is explicitly programmed. In ML systems, behavior is learned from data. This single difference creates a cascade of operational challenges that traditional DevOps cannot handle.
Traditional Software
ML System
| Dimension | Traditional Software | ML Systems |
|---|---|---|
| Source of behavior | Written by developers in code | Learned from data by algorithms |
| Output type | Deterministic (same input → same output) | Probabilistic (predictions with confidence) |
| Testing | Unit tests, integration tests, E2E | + Data validation, model evaluation, bias testing, adversarial testing |
| Bugs | Logic errors — traceable, reproducible | Data issues, drift — subtle, statistical, hard to trace |
| Degradation | Crashes loudly (exceptions, errors) | Degrades silently (accuracy slowly drops) |
| Dependencies | Libraries, APIs, databases | + Training data, features, model weights, hardware (GPUs) |
| Version control | Git for code | Git + DVC (data) + MLflow (models) + Docker (env) |
| Build time | Seconds to minutes | Minutes to days (training on large datasets) |
| Build cost | Minimal (CPU) | Expensive (GPU hours: $1-$100K+ per training run) |
| Deployment | Deploy once, update as needed | Continuous retraining + redeployment cycle |
| Monitoring | Errors, latency, uptime, memory | + Prediction quality, data drift, concept drift, fairness |
| Rollback | Revert to previous code version | Revert code + model + possibly retrain on different data |
| Compliance | Security audits, SOC2 | + Model explainability, bias audits, data privacy (GDPR) |
| Technical debt | Code complexity, outdated libs | + Data debt, configuration debt, feature debt, model debt |
This is the single most important concept in MLOps. Data drift is when the statistical properties of incoming data change compared to the data the model was trained on.
The input data distribution changes. Example: Your model was trained on customers aged 25-45. Suddenly you get a wave of customers aged 55-70.
Detection: Compare feature distributions using KS-test, PSI, or JS divergence.
The relationship between inputs and outputs changes. Example: Before COVID, "frequent traveler" predicted high spending. During COVID, it did not.
Detection: Monitor prediction accuracy against ground truth labels.
The distribution of target variable changes. Example: Fraud rate jumps from 1% to 5% due to a new attack vector.
Detection: Track target variable distribution over time.
An upstream system changes how data is generated. Example: A partner API starts sending amounts in euros instead of dollars.
Detection: Schema validation, data contracts, range checks.
# drift_detection.py — Complete working example
import pandas as pd
import numpy as np
from scipy import stats
def calculate_psi(expected, actual, buckets=10):
"""
Population Stability Index (PSI)
PSI < 0.1 → No significant drift
PSI 0.1-0.2 → Moderate drift (investigate)
PSI > 0.2 → Significant drift (retrain!)
"""
breakpoints = np.linspace(0, 100, buckets + 1)
breakpoints = np.percentile(expected, breakpoints)
expected_counts = np.histogram(expected, breakpoints)[0] / len(expected)
actual_counts = np.histogram(actual, breakpoints)[0] / len(actual)
# Avoid division by zero
expected_counts = np.clip(expected_counts, 0.001, None)
actual_counts = np.clip(actual_counts, 0.001, None)
psi = np.sum((actual_counts - expected_counts) *
np.log(actual_counts / expected_counts))
return psi
def detect_drift(training_data, production_data, features, threshold=0.15):
"""Check all features for drift"""
results = []
for feature in features:
# Method 1: PSI
psi = calculate_psi(
training_data[feature].values,
production_data[feature].values
)
# Method 2: KS Test (Kolmogorov-Smirnov)
ks_stat, ks_pvalue = stats.ks_2samp(
training_data[feature].values,
production_data[feature].values
)
# Determine drift status
drifted = psi > threshold or ks_pvalue < 0.05
results.append({
"feature": feature,
"psi": round(psi, 4),
"ks_statistic": round(ks_stat, 4),
"ks_pvalue": round(ks_pvalue, 4),
"drift_detected": drifted,
"severity": "HIGH" if psi > 0.25 else ("MEDIUM" if psi > 0.1 else "LOW")
})
return pd.DataFrame(results)
# ─── Simulate drift scenario ───
np.random.seed(42)
# Training data (what model was trained on)
train_df = pd.DataFrame({
"avg_order_value": np.random.normal(65, 20, 10000),
"days_since_order": np.random.exponential(15, 10000),
"satisfaction": np.random.normal(3.8, 0.7, 10000),
})
# Production data (3 months later — distribution shifted!)
prod_df = pd.DataFrame({
"avg_order_value": np.random.normal(52, 25, 5000), # Lower spend!
"days_since_order": np.random.exponential(25, 5000), # Less frequent!
"satisfaction": np.random.normal(3.2, 1.0, 5000), # Lower scores!
})
# Run drift detection
features = ["avg_order_value", "days_since_order", "satisfaction"]
drift_report = detect_drift(train_df, prod_df, features)
print(drift_report.to_string(index=False))
# Output:
# feature psi ks_statistic ks_pvalue drift_detected severity
# avg_order_value 0.1847 0.1532 0.0000 True MEDIUM
# days_since_order 0.2634 0.1891 0.0000 True HIGH
# satisfaction 0.1423 0.1244 0.0000 True MEDIUM
The output above shows ALL features have drifted — days_since_order has HIGH severity drift. In production, this would trigger an alert and potentially an automated retraining pipeline. Without monitoring, this model would silently give wrong predictions for weeks.
Google's famous paper "Hidden Technical Debt in Machine Learning Systems" showed that the actual ML code is a tiny fraction of a real-world ML system:
Components of a Real ML System (the ML code is the small box in the middle)
As a data engineer, you are skilled at building most of the surrounding infrastructure. The ML code (scikit-learn, PyTorch) is the easy part to learn. The hard part — the engineering around it — is what you already do.
| Debt Type | What It Looks Like | How to Fix |
|---|---|---|
| Data Debt | Undocumented data dependencies, stale features, no data contracts | Data catalogs, feature stores, data contracts |
| Pipeline Debt | Glue code connecting mismatched systems, manual steps | Standardized pipeline frameworks (ZenML, Kubeflow) |
| Configuration Debt | Hyperparameters scattered in notebooks, no version control | Config files in Git, experiment tracking (MLflow) |
| Feature Debt | Duplicate features, unused features, training-serving skew | Feature stores (Feast), feature documentation |
| Monitoring Debt | No drift detection, no alerts, no performance tracking | Evidently, Arize, Grafana dashboards |
| Reproducibility Debt | Cannot reproduce past experiments or predictions | DVC, MLflow, Docker, full pipeline versioning |
Training-serving skew is when the data processing during training differs from processing during prediction serving. This is one of the most insidious bugs in ML systems.
# ═══ TRAINING CODE (data scientist's notebook) ═══
import pandas as pd
# Feature engineering during training
train_df["order_recency"] = (
pd.to_datetime("2025-03-01") - train_df["last_order_date"]
).dt.days
# Normalize using training data statistics
mean_recency = train_df["order_recency"].mean() # = 18.5
std_recency = train_df["order_recency"].std() # = 12.3
train_df["order_recency_norm"] = (
(train_df["order_recency"] - mean_recency) / std_recency
)
# ═══ SERVING CODE (engineer's API) ═══
# BUG 1: Uses current date instead of fixed reference date
# BUG 2: Uses different mean/std (calculated from production data)
# BUG 3: Missing null handling that was in training
def predict(customer_data):
# This computes the feature DIFFERENTLY than training!
recency = (datetime.now() - customer_data["last_order_date"]).days
recency_norm = (recency - prod_mean) / prod_std # WRONG STATS!
return model.predict([[recency_norm]])
# ═══ CORRECT APPROACH: Single feature computation ═══
# Feature is computed ONCE in the feature store
# Both training and serving read from the same source
from feast import FeatureStore
store = FeatureStore(repo_path=".")
# Training: get historical features
training_features = store.get_historical_features(
entity_df=entity_df,
features=["customer_features:order_recency_norm"]
).to_df()
# Serving: get online features (SAME computation!)
serving_features = store.get_online_features(
entity_rows=[{"customer_id": "C00123"}],
features=["customer_features:order_recency_norm"]
).to_dict()
# No skew possible — same feature pipeline produces both!
Training-serving skew is fundamentally a data pipeline consistency problem. Your experience ensuring data consistency across systems is exactly what solves this. Feature stores are essentially a pattern you already understand — a single source of truth with multiple consumers.
# .github/workflows/deploy.yml (Traditional)
name: Deploy App
on: push
jobs:
deploy:
steps:
- checkout code
- run unit tests
- run integration tests
- build Docker image
- push to registry
- deploy to staging
- run smoke tests
- deploy to production
# .github/workflows/ml-deploy.yml (ML System)
name: Deploy Model
on:
push: # Code change trigger
schedule: # Scheduled retraining trigger
- cron: '0 2 * * 1'
workflow_dispatch: # Drift alert trigger
jobs:
ml-pipeline:
steps:
# ── Standard steps ──
- checkout code
- run unit tests
# ── ML-specific: Data validation ──
- validate training data schema
- check data quality (nulls, outliers, distribution)
- compare data stats vs baseline (drift check)
# ── ML-specific: Training ──
- train model on validated data
- log experiment to MLflow
# ── ML-specific: Model validation ──
- evaluate on test set
- compare vs current production model
- check for bias across demographic groups
- run adversarial tests
- GATE: only proceed if new model is better
# ── ML-specific: Registry ──
- register model in MLflow Registry
- promote to "Staging"
# ── Standard-ish steps ──
- build Docker image (with model baked in)
- deploy to staging
- run prediction smoke tests
# ── ML-specific: Canary deployment ──
- deploy to 5% of production traffic
- monitor accuracy for 2 hours
- GATE: only proceed if accuracy is stable
# ── Full deployment ──
- promote to 100% production traffic
- update monitoring dashboards
- notify team via Slack
| Metric Category | Traditional Monitoring | ML Monitoring (Additional) |
|---|---|---|
| Infrastructure | CPU, RAM, disk, network | GPU utilization, VRAM, training throughput |
| Application | Error rate, latency, throughput | Inference latency per model, batch vs real-time |
| Data | DB connections, query performance | Feature drift, data quality scores, schema changes, missing values |
| Model | N/A | Accuracy, precision, recall, F1, AUC over time |
| Predictions | N/A | Prediction distribution, confidence scores, outlier predictions |
| Fairness | N/A | Equal accuracy across demographic groups, bias metrics |
| Business | Revenue, conversion, engagement | Revenue per model prediction, false positive cost, churn reduction |
# Example: ML monitoring dashboard setup with Evidently
from evidently.report import Report
from evidently.metric_preset import (
DataDriftPreset,
DataQualityPreset,
ClassificationPreset
)
# Create a comprehensive monitoring report
monitoring_report = Report(metrics=[
DataDriftPreset(), # Feature drift
DataQualityPreset(), # Data quality checks
ClassificationPreset(), # Model performance
])
monitoring_report.run(
reference_data=training_data, # What model was trained on
current_data=this_week_data, # What we're seeing now
)
# Save as interactive HTML dashboard
monitoring_report.save_html("monitoring_dashboard.html")
# Or get programmatic results for alerting
results = monitoring_report.as_dict()
drift_detected = results["metrics"][0]["result"]["dataset_drift"]
if drift_detected:
send_slack_alert("⚠️ Data drift detected! Check dashboard.")
trigger_retraining_pipeline()
1. ML systems have 3 axes of change (code + data + model) vs 1 for traditional software.
2. Data drift is the #1 operational challenge — your data pipeline skills directly address this.
3. Training-serving skew is solved by feature stores — a pattern you already understand.
4. ML CI/CD has extra gates for data validation and model validation — you can build these.
5. ML technical debt is mostly infrastructure debt — and infrastructure is your expertise.
6. The ML code itself is ~5% of the system. The other 95% is engineering. That is you.