Chapter 02 of 12

MLOps vs Traditional Software

Why ML systems are fundamentally different from regular software — the hidden complexities of data drift, technical debt, and the three axes of change.

The Fundamental Difference

In traditional software, behavior is explicitly programmed. In ML systems, behavior is learned from data. This single difference creates a cascade of operational challenges that traditional DevOps cannot handle.

Traditional Software

Code
Rules & Logic
→
Compile / Build
→
Software
Deterministic
vs

ML System

Data
+
Code
Algorithm
→
Training
→
Model
Probabilistic

Side-by-Side Comparison (Complete)

DimensionTraditional SoftwareML Systems
Source of behaviorWritten by developers in codeLearned from data by algorithms
Output typeDeterministic (same input → same output)Probabilistic (predictions with confidence)
TestingUnit tests, integration tests, E2E+ Data validation, model evaluation, bias testing, adversarial testing
BugsLogic errors — traceable, reproducibleData issues, drift — subtle, statistical, hard to trace
DegradationCrashes loudly (exceptions, errors)Degrades silently (accuracy slowly drops)
DependenciesLibraries, APIs, databases+ Training data, features, model weights, hardware (GPUs)
Version controlGit for codeGit + DVC (data) + MLflow (models) + Docker (env)
Build timeSeconds to minutesMinutes to days (training on large datasets)
Build costMinimal (CPU)Expensive (GPU hours: $1-$100K+ per training run)
DeploymentDeploy once, update as neededContinuous retraining + redeployment cycle
MonitoringErrors, latency, uptime, memory+ Prediction quality, data drift, concept drift, fairness
RollbackRevert to previous code versionRevert code + model + possibly retrain on different data
ComplianceSecurity audits, SOC2+ Model explainability, bias audits, data privacy (GDPR)
Technical debtCode complexity, outdated libs+ Data debt, configuration debt, feature debt, model debt

Data Drift — The Silent Killer

This is the single most important concept in MLOps. Data drift is when the statistical properties of incoming data change compared to the data the model was trained on.

Types of Drift

📊 Data Drift (Feature Drift)

The input data distribution changes. Example: Your model was trained on customers aged 25-45. Suddenly you get a wave of customers aged 55-70.

Detection: Compare feature distributions using KS-test, PSI, or JS divergence.

🎯 Concept Drift

The relationship between inputs and outputs changes. Example: Before COVID, "frequent traveler" predicted high spending. During COVID, it did not.

Detection: Monitor prediction accuracy against ground truth labels.

🏷️ Label Drift

The distribution of target variable changes. Example: Fraud rate jumps from 1% to 5% due to a new attack vector.

Detection: Track target variable distribution over time.

⬆️ Upstream Data Changes

An upstream system changes how data is generated. Example: A partner API starts sending amounts in euros instead of dollars.

Detection: Schema validation, data contracts, range checks.

Drift Detection — Complete Code Example

# drift_detection.py — Complete working example
import pandas as pd
import numpy as np
from scipy import stats

def calculate_psi(expected, actual, buckets=10):
    """
    Population Stability Index (PSI)
    PSI < 0.1  → No significant drift
    PSI 0.1-0.2 → Moderate drift (investigate)
    PSI > 0.2  → Significant drift (retrain!)
    """
    breakpoints = np.linspace(0, 100, buckets + 1)
    breakpoints = np.percentile(expected, breakpoints)
    
    expected_counts = np.histogram(expected, breakpoints)[0] / len(expected)
    actual_counts = np.histogram(actual, breakpoints)[0] / len(actual)
    
    # Avoid division by zero
    expected_counts = np.clip(expected_counts, 0.001, None)
    actual_counts = np.clip(actual_counts, 0.001, None)
    
    psi = np.sum((actual_counts - expected_counts) * 
                 np.log(actual_counts / expected_counts))
    return psi

def detect_drift(training_data, production_data, features, threshold=0.15):
    """Check all features for drift"""
    results = []
    for feature in features:
        # Method 1: PSI
        psi = calculate_psi(
            training_data[feature].values,
            production_data[feature].values
        )
        
        # Method 2: KS Test (Kolmogorov-Smirnov)
        ks_stat, ks_pvalue = stats.ks_2samp(
            training_data[feature].values,
            production_data[feature].values
        )
        
        # Determine drift status
        drifted = psi > threshold or ks_pvalue < 0.05
        
        results.append({
            "feature": feature,
            "psi": round(psi, 4),
            "ks_statistic": round(ks_stat, 4),
            "ks_pvalue": round(ks_pvalue, 4),
            "drift_detected": drifted,
            "severity": "HIGH" if psi > 0.25 else ("MEDIUM" if psi > 0.1 else "LOW")
        })
    
    return pd.DataFrame(results)

# ─── Simulate drift scenario ───
np.random.seed(42)

# Training data (what model was trained on)
train_df = pd.DataFrame({
    "avg_order_value": np.random.normal(65, 20, 10000),
    "days_since_order": np.random.exponential(15, 10000),
    "satisfaction": np.random.normal(3.8, 0.7, 10000),
})

# Production data (3 months later — distribution shifted!)
prod_df = pd.DataFrame({
    "avg_order_value": np.random.normal(52, 25, 5000),   # Lower spend!
    "days_since_order": np.random.exponential(25, 5000),   # Less frequent!
    "satisfaction": np.random.normal(3.2, 1.0, 5000),     # Lower scores!
})

# Run drift detection
features = ["avg_order_value", "days_since_order", "satisfaction"]
drift_report = detect_drift(train_df, prod_df, features)
print(drift_report.to_string(index=False))

# Output:
#        feature    psi  ks_statistic  ks_pvalue  drift_detected severity
# avg_order_value  0.1847      0.1532     0.0000           True   MEDIUM
# days_since_order 0.2634      0.1891     0.0000           True     HIGH
#    satisfaction  0.1423      0.1244     0.0000           True   MEDIUM
Critical Takeaway

The output above shows ALL features have drifted — days_since_order has HIGH severity drift. In production, this would trigger an alert and potentially an automated retraining pipeline. Without monitoring, this model would silently give wrong predictions for weeks.

ML Technical Debt — The Hidden Complexity

Google's famous paper "Hidden Technical Debt in Machine Learning Systems" showed that the actual ML code is a tiny fraction of a real-world ML system:

Components of a Real ML System (the ML code is the small box in the middle)

Data Collection
Data Verification
Feature Extraction
Configuration
ML Code
(~5% of total)
Analysis Tools
Process Mgmt
Serving Infra
Monitoring

As a data engineer, you are skilled at building most of the surrounding infrastructure. The ML code (scikit-learn, PyTorch) is the easy part to learn. The hard part — the engineering around it — is what you already do.

Types of ML-Specific Technical Debt

Debt TypeWhat It Looks LikeHow to Fix
Data DebtUndocumented data dependencies, stale features, no data contractsData catalogs, feature stores, data contracts
Pipeline DebtGlue code connecting mismatched systems, manual stepsStandardized pipeline frameworks (ZenML, Kubeflow)
Configuration DebtHyperparameters scattered in notebooks, no version controlConfig files in Git, experiment tracking (MLflow)
Feature DebtDuplicate features, unused features, training-serving skewFeature stores (Feast), feature documentation
Monitoring DebtNo drift detection, no alerts, no performance trackingEvidently, Arize, Grafana dashboards
Reproducibility DebtCannot reproduce past experiments or predictionsDVC, MLflow, Docker, full pipeline versioning

Training-Serving Skew — A Common Trap

Training-serving skew is when the data processing during training differs from processing during prediction serving. This is one of the most insidious bugs in ML systems.

Example: The Bug That Costs Millions

# ═══ TRAINING CODE (data scientist's notebook) ═══
import pandas as pd

# Feature engineering during training
train_df["order_recency"] = (
    pd.to_datetime("2025-03-01") - train_df["last_order_date"]
).dt.days

# Normalize using training data statistics
mean_recency = train_df["order_recency"].mean()   # = 18.5
std_recency = train_df["order_recency"].std()     # = 12.3
train_df["order_recency_norm"] = (
    (train_df["order_recency"] - mean_recency) / std_recency
)

# ═══ SERVING CODE (engineer's API) ═══
# BUG 1: Uses current date instead of fixed reference date
# BUG 2: Uses different mean/std (calculated from production data)
# BUG 3: Missing null handling that was in training

def predict(customer_data):
    # This computes the feature DIFFERENTLY than training!
    recency = (datetime.now() - customer_data["last_order_date"]).days
    recency_norm = (recency - prod_mean) / prod_std  # WRONG STATS!
    return model.predict([[recency_norm]])

The Fix: Use a Feature Store

# ═══ CORRECT APPROACH: Single feature computation ═══
# Feature is computed ONCE in the feature store
# Both training and serving read from the same source

from feast import FeatureStore

store = FeatureStore(repo_path=".")

# Training: get historical features
training_features = store.get_historical_features(
    entity_df=entity_df,
    features=["customer_features:order_recency_norm"]
).to_df()

# Serving: get online features (SAME computation!)
serving_features = store.get_online_features(
    entity_rows=[{"customer_id": "C00123"}],
    features=["customer_features:order_recency_norm"]
).to_dict()

# No skew possible — same feature pipeline produces both!
Data Engineer Advantage

Training-serving skew is fundamentally a data pipeline consistency problem. Your experience ensuring data consistency across systems is exactly what solves this. Feature stores are essentially a pattern you already understand — a single source of truth with multiple consumers.

CI/CD: Traditional vs ML

Traditional CI/CD Pipeline

# .github/workflows/deploy.yml (Traditional)
name: Deploy App
on: push
jobs:
  deploy:
    steps:
      - checkout code
      - run unit tests
      - run integration tests
      - build Docker image
      - push to registry
      - deploy to staging
      - run smoke tests
      - deploy to production

ML CI/CD Pipeline (Notice the Extra Steps)

# .github/workflows/ml-deploy.yml (ML System)
name: Deploy Model
on:
  push:          # Code change trigger
  schedule:      # Scheduled retraining trigger
    - cron: '0 2 * * 1'
  workflow_dispatch:  # Drift alert trigger

jobs:
  ml-pipeline:
    steps:
      # ── Standard steps ──
      - checkout code
      - run unit tests
      
      # ── ML-specific: Data validation ──
      - validate training data schema
      - check data quality (nulls, outliers, distribution)
      - compare data stats vs baseline (drift check)
      
      # ── ML-specific: Training ──
      - train model on validated data
      - log experiment to MLflow
      
      # ── ML-specific: Model validation ──
      - evaluate on test set
      - compare vs current production model
      - check for bias across demographic groups
      - run adversarial tests
      - GATE: only proceed if new model is better
      
      # ── ML-specific: Registry ──
      - register model in MLflow Registry
      - promote to "Staging"
      
      # ── Standard-ish steps ──
      - build Docker image (with model baked in)
      - deploy to staging
      - run prediction smoke tests
      
      # ── ML-specific: Canary deployment ──
      - deploy to 5% of production traffic
      - monitor accuracy for 2 hours
      - GATE: only proceed if accuracy is stable
      
      # ── Full deployment ──
      - promote to 100% production traffic
      - update monitoring dashboards
      - notify team via Slack

Monitoring: Traditional vs ML

Metric CategoryTraditional MonitoringML Monitoring (Additional)
InfrastructureCPU, RAM, disk, networkGPU utilization, VRAM, training throughput
ApplicationError rate, latency, throughputInference latency per model, batch vs real-time
DataDB connections, query performanceFeature drift, data quality scores, schema changes, missing values
ModelN/AAccuracy, precision, recall, F1, AUC over time
PredictionsN/APrediction distribution, confidence scores, outlier predictions
FairnessN/AEqual accuracy across demographic groups, bias metrics
BusinessRevenue, conversion, engagementRevenue per model prediction, false positive cost, churn reduction
# Example: ML monitoring dashboard setup with Evidently
from evidently.report import Report
from evidently.metric_preset import (
    DataDriftPreset,
    DataQualityPreset, 
    ClassificationPreset
)

# Create a comprehensive monitoring report
monitoring_report = Report(metrics=[
    DataDriftPreset(),           # Feature drift
    DataQualityPreset(),         # Data quality checks
    ClassificationPreset(),      # Model performance
])

monitoring_report.run(
    reference_data=training_data,    # What model was trained on
    current_data=this_week_data,     # What we're seeing now
)

# Save as interactive HTML dashboard
monitoring_report.save_html("monitoring_dashboard.html")

# Or get programmatic results for alerting
results = monitoring_report.as_dict()
drift_detected = results["metrics"][0]["result"]["dataset_drift"]

if drift_detected:
    send_slack_alert("⚠️ Data drift detected! Check dashboard.")
    trigger_retraining_pipeline()

Key Takeaways for Your Transition

Summary

1. ML systems have 3 axes of change (code + data + model) vs 1 for traditional software.
2. Data drift is the #1 operational challenge — your data pipeline skills directly address this.
3. Training-serving skew is solved by feature stores — a pattern you already understand.
4. ML CI/CD has extra gates for data validation and model validation — you can build these.
5. ML technical debt is mostly infrastructure debt — and infrastructure is your expertise.
6. The ML code itself is ~5% of the system. The other 95% is engineering. That is you.