Site Reliability Engineer - AI & ML Infrastructure (Kubernetes, AWS & Terraform)
Core
Build and operate the hybrid infrastructure foundation for advanced AI/ML research and product development, enabling teams to train and deploy complex models at scale.
Role type
Senior Site Reliability Engineer (AI/ML Infrastructure)
Builds
Hybrid cloud and on-premise platform spanning AWS and bare metal data centers for GPU-intensive workloads
Domain
AI/ML Infrastructure, Cloud Computing, High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Kubernetes architecture and operations, Terraform (Infrastructure-as-Code), Slurm job scheduling, bare metal server management, Python/Go/Bash scripting, networking (CNI, service mesh), storage (CSI, S3)
Preferred skills
CI/CD systems, FinOps principles, Kubernetes networking/storage solutions, multi-region/hybrid cloud experience
Technologies
Kubernetes, AWS, Terraform, Slurm, Python, Go, Bash, Calico, Cilium, Ceph, Rook, GitLab CI, Jenkins, ArgoCD
Responsibilities
Architect and maintain core computing platform on AWS and on-premise; Develop and manage infrastructure using IaC with Terraform; Design and optimize AI/ML job scheduling systems integrating Slurm with Kubernetes; Provision and manage on-premise bare metal server infrastructure; Implement observability stack and automation for operational tasks; Collaborate with AI researchers to build tools accelerating development cycles
Seniority
Senior, hands-on IC