Site Reliability Engineer
Core
Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support web services and ML workloads, balancing day-to-day operations with long-term software engineering improvements.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Cloud-agnostic platform offering an abstraction layer between science and infrastructure; scalable, highly available, and fault-tolerant infrastructures for web services and ML workloads.
Domain
Cloud computing, High-Performance Computing (HPC), AI/ML infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, CI/CD, containerization, observability, Python, Go, Bash, networking, security, system administration
Preferred skills
AI/ML environment experience, HPC systems and workload managers (Slurm), modern AI-oriented solutions (Fluidstack, Coreweave, Vast)
Technologies
Kubernetes, Flux, Terraform, Docker, Prometheus, Grafana, ELK Stack, Datadog, CloudFormation, Python, Go, Bash, Slurm, Fluidstack, Coreweave, Vast
Responsibilities
Design and maintain scalable, highly available, and fault-tolerant infrastructures; operate systems and troubleshoot issues in production environments; implement and improve monitoring, alerting, and incident response systems; drive continuous improvement in infrastructure automation, deployment, and orchestration; collaborate with AI/ML researchers to develop solutions for safe and reproducible model-training experiments; document processes and procedures to ensure consistency and knowledge sharing.
Seniority
Senior, hands-on IC
