Chapter 05 of 12

Experiment Tracking

Never lose a model again - master MLflow and Weights & Biases for reproducible ML experiments.

Why Experiment Tracking?

Data scientists run hundreds of experiments. Without tracking, they lose track of what worked. Experiment tracking logs every model run: parameters, metrics, data versions, and artifacts.

Analogy

Experiment tracking is like Git for ML experiments. Every training run is a commit with full metadata.

What Gets Tracked

CategoryExamplesWhy
Parameterslearning_rate=0.05, n_estimators=200Know what settings produced each result
Metricsaccuracy=0.91, f1=0.88, auc=0.94Compare models objectively
ArtifactsModel file, plots, confusion matrixReproduce and deploy any past model
Source CodeGit commit hash, training scriptKnow what code produced the model
Data VersionDVC hash, row count, feature listFull lineage from source to prediction
EnvironmentPython version, library versionsReproduce the exact environment

MLflow Complete Setup

# Install and start
pip install mlflow
mlflow ui  # Opens at http://localhost:5000

# Production server setup
mlflow server \
    --backend-store-uri postgresql://user:pass@db:5432/mlflow \
    --default-artifact-root s3://mlflow-bucket/artifacts \
    --host 0.0.0.0 --port 5000

Track Experiments

import mlflow
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.metrics import accuracy_score, f1_score

mlflow.set_experiment("customer_churn_v2")

experiments = [
    {"n_estimators": 100, "learning_rate": 0.1, "max_depth": 3},
    {"n_estimators": 200, "learning_rate": 0.05, "max_depth": 4},
    {"n_estimators": 300, "learning_rate": 0.03, "max_depth": 5},
]

for i, params in enumerate(experiments):
    with mlflow.start_run(run_name=f"xgb_run_{i+1}"):
        mlflow.log_params(params)
        
        model = GradientBoostingClassifier(**params, random_state=42)
        model.fit(X_train, y_train)
        
        y_pred = model.predict(X_test)
        mlflow.log_metric("accuracy", accuracy_score(y_test, y_pred))
        mlflow.log_metric("f1", f1_score(y_test, y_pred))
        mlflow.sklearn.log_model(model, "model")
        
        print(f"Run {i+1}: F1={f1_score(y_test, y_pred):.3f}")

MLflow vs Weights & Biases

FeatureMLflowW&B
PricingFree (open-source)Free tier + paid
HostingSelf-hostedCloud SaaS
UI QualityGoodExcellent
CollaborationBasicExcellent
Best ForSelf-hosted, DatabricksTeams wanting best UX
Recommendation

Start with MLflow - open-source, industry standard. Your ability to deploy and manage the MLflow server is itself a valuable skill.