CareerPlanSign in

Site Reliability Engineer - AI & ML Infrastructure (Kubernetes, AWS & Terraform)

💼 Full-time🗓 2026-06-25 → 2026-09-24

Core

Build and operate the hybrid infrastructure foundation for advanced AI/ML research and product development, enabling teams to train and deploy complex models at scale.

Role type

Senior Site Reliability Engineer (AI/ML Infrastructure)

Builds

Hybrid cloud and on-premise platform spanning AWS and bare metal data centers for GPU-intensive workloads

Domain

AI/ML Infrastructure, Cloud Computing, High-Performance Computing

Deliverable

production ML models | infrastructure

Required skills

Kubernetes architecture and operations, Terraform (Infrastructure-as-Code), Slurm job scheduling, bare metal server management, Python/Go/Bash scripting, networking (CNI, service mesh), storage (CSI, S3)

Preferred skills

CI/CD systems, FinOps principles, Kubernetes networking/storage solutions, multi-region/hybrid cloud experience

Technologies

Kubernetes, AWS, Terraform, Slurm, Python, Go, Bash, Calico, Cilium, Ceph, Rook, GitLab CI, Jenkins, ArgoCD

Responsibilities

Architect and maintain core computing platform on AWS and on-premise; Develop and manage infrastructure using IaC with Terraform; Design and optimize AI/ML job scheduling systems integrating Slurm with Kubernetes; Provision and manage on-premise bare metal server infrastructure; Implement observability stack and automation for operational tasks; Collaborate with AI researchers to build tools accelerating development cycles

Seniority

Senior, hands-on IC

Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.