HPC Operations Engineering Manager
Core
Lead a team of SREs to ensure uptime, resiliency, and fault tolerance of AI model training and inference systems in hybrid cloud/on-prem environments.
Role type
Senior HPC Operations Engineering Manager
Builds
Automated deployment, incident response, scaling, and failover systems for CPU+GPU environments
Domain
High-performance computing, AI/ML infrastructure, hybrid cloud
Deliverable
infrastructure
Required skills
Kubernetes, Docker, container orchestration, Python, Go, Bash, monitoring & observability tools (Grafana, Datadog, OpenTelemetry), CI/CD pipelines, distributed systems, networking, storage, capacity planning, cost optimization
Preferred skills
Large-scale GPU cluster management, ML training/inference pipelines, HPC workload schedulers, Kubernetes operators
Technologies
Azure, AWS, GCP, Kubernetes, Docker, Grafana, Datadog, OpenTelemetry
Responsibilities
Lead team of SREs, design monitoring/alerting/logging systems, build automation for deployments and incident response, manage on-call rotations and postmortems, ensure security and compliance, partner with ML engineers to improve developer experience
Seniority
Senior, hands-on IC with people management