Study ref. INFRA-01
Infrastructure Platform Migration
Infrastructure as Code & Container Orchestration at Scale
The Situation
A mid-size e-commerce platform had grown from a single on-premises host into a sprawl of fourteen virtual machines across two cloud providers. Each server had been provisioned manually, patched ad hoc, and configured over time through SSH sessions that left no audit trail. No two hosts were identical. The deployment process was a bespoke shell script — different for every service — that a single senior engineer maintained and understood.
The consequence was chronic: a production incident in Q4 peak season revealed that three services were running different versions of the same shared library. Recovery required four hours of manual reconciliation. Horizontal scaling during peak load was impossible because new instances would diverge from production state within hours. The platform was entirely dependent on one person's institutional knowledge to recover from failure.
Constraints
- Zero-downtime migration — e-commerce cannot afford planned outages
- Multi-cloud: AWS for compute, GCP for data services — must work across both
- Team of six: one infrastructure specialist, five application engineers
- No Kubernetes experience on the team — tooling choice must be learnable
- Peak season (Q4) in 20 weeks — full migration must complete before then
Platform Evolution
Left: original state — heterogeneous VMs, no declared configuration, manual deployments. Right: target state — immutable images built from source, Terraform managing all cloud resources, Kubernetes orchestrating workloads declaratively.
The critical shift: infrastructure state lives in version control, not on individual hosts. Every pod runs from the same immutable image. Desired state is declared once; the platform converges to it — continuously.
Phase 1: Terraform Foundation (Weeks 1–4)
Before touching a single container, all existing cloud resources were imported into Terraform state. This exposed configuration drift immediately — eighteen security group rules that existed in production but not in any documentation. The import process itself served as a configuration audit.
Every subsequent cloud resource change was gated behind a terraform plan
review in CI. No engineer could modify a security group, subnet, or IAM role
outside source control. Drift detection ran nightly and opened a ticket on
any discrepancy.
Declared state flows Git Repo → Terraform Plan → Cluster; the nightly drift check evaluates each node and reconciles a fraction back to the repo boundary for re-apply — convergence enforced continuously, not just at deploy time.
Phase 2: Containerisation (Weeks 5–10)
Services were containerised in order of risk: stateless services first, stateful workers last. Each Dockerfile was written to produce deterministic images — pinned base images, no package manager operations at runtime, non-root user, read-only filesystem where possible.
A key decision was to avoid custom Helm charts initially. Plain Kubernetes
manifests in a k8s/ directory per service kept the learning
curve manageable. Complexity was deferred until the team had operational
experience with the platform.
Phase 3: Progressive Traffic Cutover (Weeks 11–16)
Traffic was shifted to Kubernetes pods using weighted DNS — 5% initially, increasing to 100% over four weeks per service. Rollback required only a DNS weight change, executable in under 90 seconds. No service required a maintenance window.
Kubernetes liveness and readiness probes replaced the previous monitoring approach, which had been manual health checks every 15 minutes. The platform now self-healed: a crashed pod was replaced within 30 seconds, automatically, without paging anyone.
Horizontal Pod Autoscaling
HPA was configured for each stateless service with CPU-based scaling and a minimum replica count of 2. Q4 peak — the original forcing function for this project — arrived eight weeks after the migration completed. Traffic peaked at 8× baseline. The platform scaled automatically; no engineer intervention was required. The incident that had taken four hours to resolve the prior year did not recur.
Deployment lead time reduction
Days of manual coordination → 4-minute CI pipeline
Snowflake servers remaining
All 14 services running from declared, versioned configuration
Mean time to recovery
Down from 4 hours — Kubernetes self-heals without manual intervention
Peak traffic handled
HPA scaled automatically through Q4 — no incidents, no pages
What I Would Do Differently
The Terraform import phase surfaced 18 undocumented security group rules — each required a decision about whether it was intentional or legacy drift. That review consumed two full weeks that were not in the original estimate. In future migrations I would schedule a dedicated configuration audit sprint before committing to an IaC import timeline.
Plain Kubernetes manifests worked well for the initial migration but became difficult to maintain across 14 services as each accumulated environment-specific patches. Adopting Kustomize overlays at week 10 — rather than week 14 — would have avoided the duplication that accumulated in the final push. The lesson: keep the tooling simple to start, but identify the abstraction threshold before you hit it, not after.
The most underestimated factor was cultural: engineers accustomed to SSH access felt a loss of control when direct host modification was prohibited. Investing in runbook documentation and a clear escalation path for "I need to change something and CI is broken" would have reduced friction in the first month of the new workflow.