Site Reliability Engineer
Core
Design, operate, and improve reliable infrastructure for large-scale AI training and inference workloads, including high-performance networks, GPU clusters, and storage.
Role type
Senior Site Reliability Engineer (AI Infrastructure)
Builds
Production-grade AI systems for natural, capable, and useful human-AI communication
Domain
Artificial Intelligence / High-Performance Computing / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Linux administration, scripting, networking (BGP, InfiniBand), cluster scheduling (Kubernetes, SLURM), distributed storage (Ceph), GPU administration, incident response, automation, capacity planning
Preferred skills
NVIDIA GPU/CUDA/NCCL expertise, high-performance interconnects (RDMA, RoCE), Terraform/Ansible, observability (Prometheus, Grafana), bare-metal automation, cloud infrastructure (AWS, GCP, Azure)
Technologies
Kubernetes, SLURM, MAAS, Ceph, InfiniBand, CUDA, Prometheus, Grafana, Terraform, Ansible, AWS, GCP, Azure
Responsibilities
Design and operate infrastructure for AI training/inference; automate operational workflows; build monitoring and incident-response practices; diagnose performance and reliability issues; partner with ML/research teams; improve provisioning and deployment automation; plan cluster growth and lifecycle management
Seniority
Senior, hands-on IC