ML Model Versioning and Registry Patterns for Production

How to implement model versioning and registry patterns that enable reproducible deployments, safe rollbacks, and audit trails for production ML systems.

#machine-learning#model-registry#mlops#versioning
Cover image for the article: ML Model Versioning and Registry Patterns for Production

Running machine learning models in production without proper versioning is like deploying code without Git. You eventually lose track of what's running where, rollbacks become impossible, and debugging regressions turns into guesswork. After managing 40+ production models across multiple teams, I've developed patterns for model versioning and registry management that bring the same rigor to ML that we expect from software engineering.

The Problem: Model Chaos at Scale

Most teams start by saving models as timestamped files in S3. This works until your third model, at which point you discover:

  • Nobody knows which model version is serving traffic in production
  • Rolling back a bad model requires someone to remember the exact artifact path
  • There's no audit trail for who promoted a model or why
  • Reproducing a specific model version requires detective work across training scripts, data snapshots, and hyperparameters

The cost of this chaos compounds. In one incident, a fraud detection model was accidentally rolled back two versions during a deployment, resulting in 72 hours of elevated false negatives before anyone noticed.

Architecture: A Three-Layer Registry

The registry architecture I've found most effective separates concerns into three layers: artifact storage, metadata management, and lifecycle governance.

Model Registry Architecture

Layer 1: Artifact Storage

Model artifacts live in immutable, content-addressed storage. Every artifact gets a SHA-256 hash, and that hash becomes its canonical identifier. This guarantees that a specific version always refers to exactly the same binary.

import hashlib
import json
from dataclasses import dataclass, field
from datetime import datetime
from typing import Optional
import boto3

@dataclass
class ModelArtifact:
    model_id: str
    version: str
    artifact_hash: str
    artifact_uri: str
    framework: str
    created_at: datetime = field(default_factory=datetime.utcnow)
    metadata: dict = field(default_factory=dict)

class ArtifactStore:
    def __init__(self, bucket: str, prefix: str = "models"):
        self.s3 = boto3.client("s3")
        self.bucket = bucket
        self.prefix = prefix

    def register_artifact(self, model_bytes: bytes, model_id: str, 
                          framework: str, metadata: dict) -> ModelArtifact:
        artifact_hash = hashlib.sha256(model_bytes).hexdigest()
        version = f"v{datetime.utcnow().strftime('%Y%m%d%H%M%S')}-{artifact_hash[:8]}"
        artifact_key = f"{self.prefix}/{model_id}/{version}/model.bin"

        self.s3.put_object(
            Bucket=self.bucket,
            Key=artifact_key,
            Body=model_bytes,
            Metadata={"content-hash": artifact_hash},
            ContentType="application/octet-stream",
        )

        return ModelArtifact(
            model_id=model_id,
            version=version,
            artifact_hash=artifact_hash,
            artifact_uri=f"s3://{self.bucket}/{artifact_key}",
            framework=framework,
            metadata=metadata,
        )

    def verify_integrity(self, artifact: ModelArtifact) -> bool:
        response = self.s3.get_object(
            Bucket=self.bucket,
            Key=artifact.artifact_uri.replace(f"s3://{self.bucket}/", ""),
        )
        actual_hash = hashlib.sha256(response["Body"].read()).hexdigest()
        return actual_hash == artifact.artifact_hash

Layer 2: Metadata Management

The metadata layer captures everything needed to reproduce and understand a model version. This includes training lineage, evaluation metrics, feature schemas, and deployment configuration.

interface ModelVersion {
  modelId: string;
  version: string;
  artifactHash: string;
  artifactUri: string;
  stage: 'development' | 'staging' | 'production' | 'archived';
  trainingLineage: {
    datasetVersion: string;
    datasetHash: string;
    featureSchema: Record<string, string>;
    hyperparameters: Record<string, number | string | boolean>;
    trainingJobId: string;
    trainingDuration: number;
    framework: string;
    frameworkVersion: string;
  };
  evaluation: {
    metrics: Record<string, number>;
    evaluationDataset: string;
    evaluatedAt: string;
    thresholds: Record<string, { min?: number; max?: number }>;
  };
  deployment: {
    servingConfig: {
      batchSize: number;
      maxLatencyMs: number;
      instanceType: string;
      minReplicas: number;
      maxReplicas: number;
    };
    canaryPercentage: number;
    rollbackVersion?: string;
  };
  governance: {
    promotedBy: string;
    promotedAt: string;
    approvedBy: string[];
    reason: string;
    reviewUrl?: string;
  };
}

class ModelRegistry {
  private db: DynamoDB;

  async promoteModel(
    modelId: string,
    version: string,
    targetStage: ModelVersion['stage'],
    promoter: string,
    reason: string
  ): Promise<void> {
    const model = await this.getVersion(modelId, version);

    if (targetStage === 'production') {
      await this.validatePromotionGates(model);
    }

    const currentProd = await this.getCurrentProduction(modelId);

    await this.db.transactWrite([
      { update: { modelId, version, stage: targetStage, 
                  governance: { promotedBy: promoter, promotedAt: new Date().toISOString(), reason } } },
      ...(currentProd ? [{ update: { modelId, version: currentProd.version, stage: 'archived' as const } }] : []),
      { put: { type: 'audit_log', modelId, version, action: 'promote', 
               targetStage, promoter, reason, timestamp: new Date().toISOString() } },
    ]);
  }

  private async validatePromotionGates(model: ModelVersion): Promise<void> {
    const { metrics, thresholds } = model.evaluation;
    for (const [metric, threshold] of Object.entries(thresholds)) {
      const value = metrics[metric];
      if (threshold.min !== undefined && value < threshold.min) {
        throw new Error(`Metric ${metric}=${value} below minimum ${threshold.min}`);
      }
      if (threshold.max !== undefined && value > threshold.max) {
        throw new Error(`Metric ${metric}=${value} above maximum ${threshold.max}`);
      }
    }
  }
}

Layer 3: Lifecycle Governance

The governance layer enforces promotion policies, manages approval workflows, and maintains the audit trail. The key principle is that no model reaches production without passing through automated quality gates and human approval.

Versioning Strategy: Semantic + Content Addressing

I use a hybrid versioning scheme that combines semantic meaning with content addressing:

  • Major version: Changes to model architecture or feature schema
  • Minor version: Retraining on new data with same architecture
  • Patch version: Configuration changes, serving optimization
  • Content hash suffix: Guarantees artifact identity

A version string looks like v2.4.1-a3f8c2b1, where the semantic part communicates intent and the hash suffix guarantees exactness.

Deployment Patterns

Blue-Green Model Deployment

Maintain two production slots. The "blue" slot serves current traffic while "green" receives the new model. Traffic switches atomically after validation passes.

Shadow Mode Testing

New model versions run in shadow mode first, receiving production traffic but not serving responses. This catches issues that only manifest with real-world data distributions.

Canary Promotion

Route a small percentage (typically 5-10%) of traffic to the new version. Monitor key metrics for degradation before increasing traffic. Automated rollback triggers if error rates or latency exceed thresholds.

Benchmarks: Registry Performance

After implementing this pattern across our platform:

MetricBeforeAfter
Time to identify running model version15-30 min< 1 sec
Rollback time (model swap)20-45 min90 sec
Reproducibility rate~40%99.8%
Promotion with audit trail0%100%
Failed deployment recoveryManual, hoursAutomated, < 3 min

Lessons from Production

Immutability is non-negotiable. Once an artifact is registered, it must never be modified. If you need to change something, create a new version. This seems obvious but requires discipline when someone wants to "just fix the config."

Version everything together. The model, its feature schema, serving configuration, and preprocessing code are a unit. Versioning the model alone leads to subtle bugs when serving code evolves independently.

Automate governance but keep humans in the loop. Automated quality gates catch most issues, but production promotion for critical models should still require human approval. The system should make that approval easy by presenting all relevant metrics and comparisons.

Build for rollback from day one. Every deployment should know its predecessor. Rollback should be a single operation, not an emergency scramble. Test your rollback procedure monthly.

Conclusion

Model versioning and registry management is infrastructure that pays for itself the first time you need to answer "what changed?" after a production regression. The patterns here — content-addressed storage, rich metadata, lifecycle governance — bring ML deployments to the same maturity level we expect from software deployments. Start with immutable artifacts and an audit trail, then add governance layers as your model count grows.

Comments

    No comments yet. Be the first to share your thoughts.