Migrating production workloads to Kubernetes is one of those projects that sounds simple in architecture review meetings and becomes incredibly complex once you start executing. Over the past three years, we have migrated 47 production systems to Kubernetes for clients ranging from Series A startups to Fortune 500 enterprises. Here is what we have learned about doing it without waking up your on-call engineers at 3AM.
The most critical lesson is this: never do a big-bang migration. We use a progressive migration strategy that starts with non-critical workloads (internal tools, staging environments, batch processing jobs) and gradually moves to customer-facing services. This approach lets your team build Kubernetes operational expertise on low-risk systems before tackling the services that directly impact revenue.
Traffic shifting is where the magic happens. We use a service mesh (typically Istio or Linkerd) to gradually shift traffic from the legacy infrastructure to Kubernetes. Starting at 1% and monitoring error rates, latency percentiles, and resource utilization at each step. The key metrics we watch are P99 latency, error rate, and pod restart counts. If any metric degrades beyond our predefined thresholds, traffic automatically shifts back.
Database migrations during a Kubernetes move deserve their own strategy. We use a dual-write pattern where both the old and new systems write to the same database during the transition period. For systems requiring a database migration (say, moving from a VM-hosted PostgreSQL to a managed service), we set up logical replication first, validate data consistency for at least two weeks, then cut over the read traffic before the write traffic.
Observability must be in place before the migration begins, not after. We deploy a full monitoring stack (Prometheus, Grafana, distributed tracing with Jaeger, and structured logging with Loki) into the Kubernetes cluster and run it alongside the legacy monitoring for at least a month. This overlap period lets you build confidence in the new monitoring and ensures you won't lose visibility during the transition.
The operational runbooks matter as much as the technical implementation. We create detailed runbooks for every failure scenario we can imagine: pod crashes, node failures, network partitions, certificate expirations, resource exhaustion. Each runbook is tested in a chaos engineering session before the production migration begins. Teams that skip this step invariably learn the hard way during their first production incident on Kubernetes.
Topics