Why Big Bang Rewrites Fail
The monolith handled $50M ARR. 2.3M lines of Java. Zero tests. The business couldn't pause for 18 months. We chose the Strangler Fig: incrementally extract, route traffic, verify, repeat.
Step 1: The Facade Layer
Put an API gateway in front of the monolith. Every external call hits the gateway first. New services register routes. The gateway routes by path — old paths to monolith, new paths to services.
Step 2: Identify the First Seam
We picked "User Notifications" — low risk, well-bounded, high change rate. Extracted to a Go service in 3 weeks. Ran both in parallel with shadow traffic for 2 weeks before cutting over.
Step 3: Contract Testing as the Safety Net
Pact tests between gateway and each service. Monolith's consumer contracts generated from existing traffic. No deploy without contract validation. This caught 12 breaking changes in year one.
Step 4: Data Ownership Transfer
The hardest part. We used dual-write with CDC (Debezium) to sync the monolith's tables to the new service's database. Ran reconciliation jobs nightly. Cut over reads first, then writes.
Step 5: Organizational Alignment
Each extracted service got a dedicated team with full ownership: code, infra, on-call, product decisions. No more "ticket to the platform team." This cultural shift mattered more than the technical pattern.
18 Months Later
- 12 services extracted
- Monolith down to 800K lines (mostly stable core)
- Deploy frequency: monthly → daily
- MTTR: 4 hours → 15 minutes
- Zero major incidents from extractions