Tommy Cho ML Infrastructure

Senior Software Engineer, ML Infrastructure · CoinTracker (YC W18)

I make complex infrastructure behave.

Hi, my name is Tommy. I build the platforms behind production systems, reducing cloud spend, improving latency, and making distributed infrastructure easier to operate under pressure.

Current focus
Production ML infrastructure
Measured impact
$658K+ annualized savings
Based in
Toronto, Canada
ML Infrastructure Kubernetes Istio Observability Cloud Economics Distributed Systems

01 / Selected impact

Proof over
buzzwords.

A few examples of turning infrastructure problems into measurable business and engineering outcomes.

Case study 02 Service mesh

Simplified cross-Atlantic service discovery.

Led the rearchitecture from NGINX reverse proxies behind AWS NLBs to native Istio ingress and egress gateways, reducing infrastructure and potential failure points.

−15%P99 latency
−97%xDS bandwidth
Case study 03 ML × observability

Used local models to make incidents legible faster.

Deployed GPT-OSS with vLLM across a multi-cluster Kubernetes environment to surface LLM-driven observability insights and accelerate production diagnosis.

−90%Mean time to recovery
  • vLLM
  • Multi-cluster Kubernetes
  • Production observability

02 / Operating principles

How I think when systems get expensive or weird.

  1. 01

    Cost is an architecture signal.

    A large bill often exposes duplication, poor boundaries, idle capacity, or the wrong managed-service tradeoff.

  2. 02

    Make failure visible before making it rare.

    You cannot improve a system you cannot explain during an incident.

  3. 03

    Build paved roads, not permission queues.

    Federated delivery platforms and safe rollout primitives let developers move quickly without bypassing controls.

  4. 04

    Optimize the whole operating loop.

    Deployment, telemetry, diagnosis, rollback, and learning are one reliability system—not separate tool purchases.

03 / Current focus

Building the infrastructure for inference.

At CoinTracker, I’m applying the same systems mindset to ML infrastructure, where model performance, evaluation, and training all become part of the same engineering problem.

Building production ML infrastructure and the workflows that support it.

04 / CV snapshot

Experience, compressed to signal.

Senior Software Engineer, ML Infrastructure

  • Building production ML infrastructure and the workflows that support it.

Site Reliability Engineer II

  • Delivered $108K per year in cloud savings through Kubernetes optimization, Redis-engine migration, and automated cleanup of detached EC2 instances.
  • Reduced cross-Atlantic service-discovery P99 latency by 15% by replacing NGINX proxies and AWS NLBs with native Istio gateways.
  • Reduced Istio xDS configuration-update bandwidth by up to 97% by scoping egress host changes to affected workloads.
  • Improved production reliability through post-incident root-cause analysis and corrective actions.

Site Reliability Engineer II

  • Saved more than $250K annually by consolidating Confluent Kafka clusters to one per environment.
  • Reduced compute costs by up to 75% through Kubernetes rightsizing and Goldilocks-driven node-pool tuning.
  • Cut $300K in observability spend by migrating Grafana, Prometheus, Loki, and Tempo to a self-hosted stack on GKE.
  • Improved scalability by 50% through proactive autoscaling with KEDA and custom metrics.
  • Reduced MTTR by 90% by deploying GPT-OSS with vLLM for observability insights across multiple Kubernetes clusters.
  • Built developer delivery capabilities with federated ArgoCD, Argo Rollouts, Skaffold, and Cloud Build.
  • Maintained Flow blockchain nodes and safely executed protocol upgrades, configuration changes, and restarts.

DevOps, software engineering, and systems support

  • Worked across DevOps, software engineering, and technical operations at TrustFlight and the University of British Columbia. Built infrastructure delivery workflows with Terraform, ArgoCD, and Skaffold; centralized cloud security logs in Splunk; and helped migrate a production Drupal platform while improving page-load performance by 35%.