Senior Site Reliability Engineer, Production Engineering
Core
Senior SRE responsible for maintaining high availability, reliability, and security of large-scale cloud and infrastructure services, specifically focusing on production Kubernetes environments.
Role type
Senior Site Reliability Engineer (Production Engineering)
Builds
Large-scale Kubernetes clusters, high-performance computing environments, and global 24/7 reliability operations
Domain
Cloud infrastructure, High-Performance Computing (HPC), Kubernetes, Bare-metal systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes administration, Linux systems administration, Networking (DNS, DHCP, IP tables, routing, firewalls), Incident management, Observability, Automation/Scripting, CI/CD pipelines, Troubleshooting complex infrastructure issues
Preferred skills
GPU/DPU hardware knowledge, Python/Golang/Rust programming, Architecture of large-scale Kubernetes environments, SLURM experience
Technologies
Kubernetes, SLURM, Jenkins, ArgoCD, Linux, Python, Golang, Rust
Responsibilities
Administer and maintain large-scale Kubernetes clusters and infrastructure; Automate operational processes to reduce manual tasks; Proactively detect and respond to production incidents using monitoring and observability; Lead incident management calls and coordinate resolution of critical issues; Analyze logs and metrics to troubleshoot complex problems; Contribute to the architecture and deployment of large-scale Kubernetes environments
Seniority
Senior, hands-on IC
