Site Reliability Engineer II
Core
Ensure the reliability, scalability, and security of AI infrastructure supporting HPC & AI workloads.
Role type
Senior Site Reliability Engineer (AI/High-Performance Computing)
Builds
AI infrastructure, containerized environments, automation tools for deployment and monitoring
Domain
Artificial Intelligence, High-Performance Computing, Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
incident management, root cause analysis, performance optimization, infrastructure automation, cloud platform management, distributed systems design, GPU/InfiniBand management, container orchestration, predictive analysis
Preferred skills
technical leadership, cross-functional collaboration, innovation in best practices
Technologies
Kubernetes, Docker, Azure, GPUs, InfiniBand
Responsibilities
Lead incident response and root cause analysis to minimize downtime. Identify and resolve bottlenecks in compute, storage, networking, and specialized hardware. Develop and maintain automation tools for deployment, monitoring, and management of AI infrastructure. Provide technical guidance in cloud and AI infrastructure technologies.
Seniority
Senior, hands-on IC with technical leadership