Operations Engineer, Fleet Reliability
Core
Provisioning, management, and uptime of CoreWeave's fleet of server nodes and supercomputing clusters.
Role type
Operations Engineer, Fleet Reliability
Builds
High-performance supercomputing clusters with state-of-the-art GPUs for AI workloads
Domain
Cloud Infrastructure / HPC / AI Compute
Deliverable
infrastructure
Required skills
Linux system administration, hardware troubleshooting, software troubleshooting, scripting (bash, python, powershell)
Preferred skills
Observability platforms (Grafana, Prometheus), Kubernetes administration, HPC/GPU workload administration, data center environment experience
Technologies
Linux, Kubernetes, Grafana, Prometheus, bash, python, powershell
Responsibilities
Configure and maintain large-scale high-performance supercomputing clusters; Troubleshoot hardware and software issues; Monitor and analyze system performance; Create and maintain documentation of team processes; Participate in oncall rotations
Seniority
Mid-level, hands-on IC