Sr. Computer Scientist - Infra Engineering
Core
Architecting and scaling Kubernetes-native infrastructure for massive GPU training workloads and AI model serving.
Role type
Senior Infrastructure Engineer (GPU/ML Platform)
Builds
Cloud infrastructure, Kubernetes operators, observability stacks, and developer tooling for ML training pipelines.
Domain
Cloud Infrastructure / Machine Learning / GPU Computing
Deliverable
infrastructure
Required skills
Kubernetes internals, Go/Python/Rust, Terraform/Pulumi, AWS, Networking (TCP/IP, CNI, BGP), Observability (Prometheus, OpenTelemetry), Distributed Systems Design
Preferred skills
Azure/GCP, eBPF, Chaos Engineering, Open Source contributions
Technologies
Kubernetes, AWS, Terraform, Pulumi, Go, Python, Rust, Prometheus, Grafana, OpenTelemetry, Jaeger, Loki, Cilium, Istio, Helm, ArgoCD
Responsibilities
Design and evolve Kubernetes-native infrastructure for distributed GPU training; Lead cross-geo initiatives and multi-team projects; Define and ship cloud infrastructure via IaC; Build automation, operators, and platform services; Lead incident response and reliability reviews; Debug complex cluster networking issues; Mentor and grow the team.
Seniority
Senior, hands-on IC with leadership responsibilities