CareerPlanSign in

Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

USA | Remote💼 Full-time🗓 2026-07-03 → 2026-09-25

Core

Build and operate the hybrid infrastructure foundation for advanced AI/ML research and product development, enabling teams to train and deploy complex models at scale.

Role type

Senior Platform Engineer (AI/ML Infrastructure)

Builds

Self-service hybrid computing platform spanning AWS and bare metal data centers

Domain

Cloud Infrastructure & AI/ML Systems

Deliverable

infrastructure

Required skills

Kubernetes architecture, Terraform, Python, Go, Bash, CI/CD systems, observability stack design, networking (CNI, service mesh), storage (CSI, S3)

Preferred skills

Slurm job scheduling, bare metal server management, FinOps, Kubernetes networking (Calico, Cilium), storage (Ceph, Rook), multi-region/hybrid cloud experience

Technologies

Kubernetes, Terraform, AWS, Slurm, GitLab CI, Jenkins, ArgoCD, Calico, Cilium, Ceph, Rook, S3

Responsibilities

Architect and maintain core computing platform on AWS and on-premise; Develop and manage infrastructure using Infrastructure-as-Code; Design and optimize AI/ML job scheduling systems integrating Slurm with Kubernetes; Provision and maintain on-premise bare metal GPU infrastructure; Implement networking and storage solutions; Develop observability stack and automation for operational tasks; Collaborate with AI researchers to build accelerating workflows.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.