Back to Blog

Kubernetes Cost Optimization: Cutting 40% Without Touching Performance

The $200K/Month Wake-Up Call

Our three EKS clusters were burning $200K/month. CPU utilization averaged 12%. Memory: 18%. We weren't even close to capacity — we were just over-provisioned everywhere.

Phase 1: Rightsizing (Week 1-2)

Used VPA in recommendation mode for two weeks, then applied. Cut 35% of CPU requests, 28% of memory. Zero incidents. Most teams had copied resource limits from a template three years ago.

Phase 2: Spot Instance Strategy (Week 3-4)

Moved stateless, fault-tolerant workloads (batch jobs, CI runners, dev environments) to Spot. Used Karpenter with capacity-type preferences. Saved 65% on those node groups. Critical path stays on On-Demand.

Phase 3: Bin Packing & Pod Topology (Week 5-6)

Enabled pod topology spread constraints. Used Descheduler to evict poorly placed pods. Consolidated from 180 nodes to 110. The remaining nodes run hotter (65% CPU) but safely.

Phase 4: Workload-Aware Scheduling (Ongoing)

ML training jobs get GPU nodes with time-limited leases. Web services get burstable classes with strict limits. Batch jobs get best-effort. The scheduler now understands what it's placing, not just where.

Results