Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Core
Build and operate the hybrid infrastructure foundation for advanced AI/ML research and product development, enabling teams to train and deploy complex models at scale.
Role type
Senior Platform Engineer (AI/ML Infrastructure)
Builds
Self-service hybrid computing platform spanning AWS and bare metal data centers
Domain
Cloud Infrastructure & AI/ML Systems
Deliverable
infrastructure
Required skills
Kubernetes architecture, Terraform, Python, Go, Bash, CI/CD systems, observability stack design, networking (CNI, service mesh), storage (CSI, S3)
Preferred skills
Slurm job scheduling, bare metal server management, FinOps, Kubernetes networking (Calico, Cilium), storage (Ceph, Rook), multi-region/hybrid cloud experience
Technologies
Kubernetes, Terraform, AWS, Slurm, GitLab CI, Jenkins, ArgoCD, Calico, Cilium, Ceph, Rook, S3
Responsibilities
Architect and maintain core computing platform on AWS and on-premise; Develop and manage infrastructure using Infrastructure-as-Code; Design and optimize AI/ML job scheduling systems integrating Slurm with Kubernetes; Provision and maintain on-premise bare metal GPU infrastructure; Implement networking and storage solutions; Develop observability stack and automation for operational tasks; Collaborate with AI researchers to build accelerating workflows.
Seniority
Senior, hands-on IC
