Study ref. BE-01
Backend API Redesign
Hexagonal Architecture & CQRS at Scale
The Situation
A payments platform had grown from an internal tool into a customer-facing API serving hundreds of third-party integrations. The original codebase was a single Express.js monolith: business logic interleaved with SQL queries, HTTP handlers directly instantiating database connections, and no meaningful test coverage below the integration layer.
Every release required a 2-hour manual regression suite. Onboarding a new engineer to a domain area averaged six weeks. When traffic doubled in Q3, the team spent three weekends firefighting cascading timeouts that originated in untestable, deeply nested query logic.
Constraints
- Zero-downtime migration — 94 external integrations cannot break
- Existing database schema must be preserved (migration risk too high)
- Team of five: two seniors, three mid-level engineers unfamiliar with DDD
- No greenfield rewrite — iterative strangling of the monolith
System Evolution
Left: original layered monolith with tight coupling between transport, business logic, and persistence. Right: hexagonal target state with isolated domain core and read/write separation via CQRS.
The critical insight: the domain core imports nothing from infrastructure. Dependency arrows always point inward. Adapters are swappable — the domain logic is unchanged whether persistence is Postgres, Redis, or an in-memory test double.
Migration Strategy: Strangler Fig
Rather than a big-bang rewrite, new domain bounded contexts were built alongside the existing monolith. The API gateway layer was introduced to route requests — legacy endpoints remained on the monolith while new command and query handlers were stood up in the hexagonal core.
Over 14 weeks, each domain area (payments, accounts, notifications) was migrated in sequence. Feature flags controlled traffic shifting per endpoint, allowing rollback within minutes if issues surfaced.
Traffic at the API gateway shifts gradually from the legacy monolith to the hexagonal core as each bounded context is migrated — the strangler fig pattern in motion.
CQRS Split Decision
The read/write ratio was 9:1 (reads dominated). Separating the query path allowed introduction of a read replica without touching the write path. Command handlers became lean: accept a validated command object, execute domain logic, persist via the outbound port, publish a domain event.
Query handlers bypassed the domain model entirely for simple reads — thin projections directly from the read replica. This eliminated N+1 queries that had caused the timeout cascades.
Testing Architecture
The domain core — containing all business rules — is pure TypeScript with zero infrastructure imports. The full suite of domain unit tests runs in under 800 ms. No database. No HTTP. No test containers.
Adapters are tested at the integration layer with a real database in CI. Port contracts are verified independently. The regression suite shrank from a 2-hour manual process to a 12-minute automated pipeline.
Team Enablement
The hexagonal boundary provided a clear mental model for engineers unfamiliar with DDD. A simple rule — "if you're writing a new feature, start in the domain core with a failing test; only touch adapters once the logic is correct in isolation" — was enough to guide the team through the migration with high confidence.
Regression time reduction
2-hour manual suite → 12-minute automated pipeline
Throughput under sustained load
Read replica offloaded query pressure from the write path
Integration breakages
94 external integrations migrated without incidents
New engineer time-to-productivity
Down from 6 weeks; clear hexagonal boundary as mental model
What I Would Do Differently
The CQRS split introduced eventual consistency in the read projection layer — a trade-off that was not communicated clearly to the product team initially. Two edge cases in notifications relied on read-after-write consistency. Both required compensating logic. A domain event log (the write side publishing events the query side consumes) would have made this explicit and testable from day one rather than discovered in staging.
The strangler fig approach worked well technically, but the feature-flag routing layer added operational overhead for three months after the migration completed. A harder cutover date — with a one-week rollback window rather than indefinite routing — would have been cleaner.