Case study · Plate v-A · Infrastructure Architecture

Study ref. INFRA-01

Infrastructure Platform Migration

Infrastructure as Code & Container Orchestration at Scale

Context E-commerce platform, 14 production services
Scale ~3M daily active users, peak 8× baseline
Duration 16 weeks (rolling cutover)
Role Lead Solution Architect
I — Context & Problem Statement

The Situation

A mid-size e-commerce platform had grown from a single on-premises host into a sprawl of fourteen virtual machines across two cloud providers. Each server had been provisioned manually, patched ad hoc, and configured over time through SSH sessions that left no audit trail. No two hosts were identical. The deployment process was a bespoke shell script — different for every service — that a single senior engineer maintained and understood.

The consequence was chronic: a production incident in Q4 peak season revealed that three services were running different versions of the same shared library. Recovery required four hours of manual reconciliation. Horizontal scaling during peak load was impossible because new instances would diverge from production state within hours. The platform was entirely dependent on one person's institutional knowledge to recover from failure.

Constraints

  • Zero-downtime migration — e-commerce cannot afford planned outages
  • Multi-cloud: AWS for compute, GCP for data services — must work across both
  • Team of six: one infrastructure specialist, five application engineers
  • No Kubernetes experience on the team — tooling choice must be learnable
  • Peak season (Q4) in 20 weeks — full migration must complete before then
II — Architecture Diagram · Before & After

Platform Evolution

Left: original state — heterogeneous VMs, no declared configuration, manual deployments. Right: target state — immutable images built from source, Terraform managing all cloud resources, Kubernetes orchestrating workloads declaratively.

Before — Snowflake VM Estate
MANUAL VM ESTATE VM-A node 14.8 · custom nginx 1.18 VM-B node 16.2 · patched nginx 1.22 VM-C node 14.15 · drift apache 2.4 ↓ SSH DEPLOY (MANUAL) ↓ SHARED DATABASE single host, no replica NO TWO VMs IDENTICAL UNDOCUMENTED CONFIGURATION SINGLE POINT OF FAILURE
After — Declarative Platform
TERRAFORM cloud resources declared in source CI PIPELINE immutable image build + push KUBERNETES CLUSTER POD image:v2.4.1 POD image:v2.4.1 POD image:v2.4.1 IDENTICAL · IMMUTABLE · REPRODUCIBLE MANAGED DB RDS · declared in Terraform

The critical shift: infrastructure state lives in version control, not on individual hosts. Every pod runs from the same immutable image. Desired state is declared once; the platform converges to it — continuously.

III — Approach & Key Decisions

Phase 1: Terraform Foundation (Weeks 1–4)

Before touching a single container, all existing cloud resources were imported into Terraform state. This exposed configuration drift immediately — eighteen security group rules that existed in production but not in any documentation. The import process itself served as a configuration audit.

Every subsequent cloud resource change was gated behind a terraform plan review in CI. No engineer could modify a security group, subnet, or IAM role outside source control. Drift detection ran nightly and opened a ticket on any discrepancy.

Declared state flows Git Repo → Terraform Plan → Cluster; the nightly drift check evaluates each node and reconciles a fraction back to the repo boundary for re-apply — convergence enforced continuously, not just at deploy time.

Phase 2: Containerisation (Weeks 5–10)

Services were containerised in order of risk: stateless services first, stateful workers last. Each Dockerfile was written to produce deterministic images — pinned base images, no package manager operations at runtime, non-root user, read-only filesystem where possible.

A key decision was to avoid custom Helm charts initially. Plain Kubernetes manifests in a k8s/ directory per service kept the learning curve manageable. Complexity was deferred until the team had operational experience with the platform.

Phase 3: Progressive Traffic Cutover (Weeks 11–16)

Traffic was shifted to Kubernetes pods using weighted DNS — 5% initially, increasing to 100% over four weeks per service. Rollback required only a DNS weight change, executable in under 90 seconds. No service required a maintenance window.

Kubernetes liveness and readiness probes replaced the previous monitoring approach, which had been manual health checks every 15 minutes. The platform now self-healed: a crashed pod was replaced within 30 seconds, automatically, without paging anyone.

Horizontal Pod Autoscaling

HPA was configured for each stateless service with CPU-based scaling and a minimum replica count of 2. Q4 peak — the original forcing function for this project — arrived eight weeks after the migration completed. Traffic peaked at 8× baseline. The platform scaled automatically; no engineer intervention was required. The incident that had taken four hours to resolve the prior year did not recur.

IV — Results & Reflection
96%

Deployment lead time reduction

Days of manual coordination → 4-minute CI pipeline

0

Snowflake servers remaining

All 14 services running from declared, versioned configuration

4 min

Mean time to recovery

Down from 4 hours — Kubernetes self-heals without manual intervention

8×

Peak traffic handled

HPA scaled automatically through Q4 — no incidents, no pages

What I Would Do Differently

The Terraform import phase surfaced 18 undocumented security group rules — each required a decision about whether it was intentional or legacy drift. That review consumed two full weeks that were not in the original estimate. In future migrations I would schedule a dedicated configuration audit sprint before committing to an IaC import timeline.

Plain Kubernetes manifests worked well for the initial migration but became difficult to maintain across 14 services as each accumulated environment-specific patches. Adopting Kustomize overlays at week 10 — rather than week 14 — would have avoided the duplication that accumulated in the final push. The lesson: keep the tooling simple to start, but identify the abstraction threshold before you hit it, not after.

The most underestimated factor was cultural: engineers accustomed to SSH access felt a loss of control when direct host modification was prohibited. Investing in runbook documentation and a clear escalation path for "I need to change something and CI is broken" would have reduced friction in the first month of the new workflow.