Senior Site Reliability Engineer
Core
Ensure reliability, scalability, and security of AI infrastructure supporting HPC & AI workloads, leading incident response and performance optimization.
Role type
Senior Site Reliability Engineer (AI/Cloud Infrastructure)
Builds
AI infrastructure services, containerized environments, and automation tools for deployment and monitoring.
Domain
Artificial Intelligence, High-Performance Computing, Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Incident management, root cause analysis, performance optimization, infrastructure automation, Kubernetes, Docker, cloud platforms, distributed systems, GPU management, InfiniBand, predictive analysis
Preferred skills
Publications, certifications in cloud or AI infrastructure
Responsibilities
Lead incident response and root cause analysis, identify and resolve bottlenecks in compute/storage/networking, develop automation tools for deployment and monitoring, provide technical guidance on cloud and AI technologies, advocate for service excellence, research emerging AI infrastructure technologies
Seniority
Senior, hands-on IC with technical leadership