Removed more than $658K in annual infrastructure spend.
This was not one optimization. It was a repeated pattern of finding expensive system seams, simplifying them, and leaving behind safer defaults.
- $300K Moved Grafana, Prometheus, Loki, and Tempo to a self-hosted stack on GKE.
- $250K+ Consolidated Confluent Kafka clusters to one per environment.
- 75% Reduced compute cost through Kubernetes rightsizing and node-pool tuning.
- $108K Cut cloud spend at PagerDuty through K8s optimization, Redis migration, and EC2 cleanup.
Cost reduction without treating reliability as collateral damage.