CareerPlanGet AI match score →

Operations Engineer, Fleet Reliability

Bellevue, WA💼 Full-time💰 $83,000–$83,000🗓 2026-07-17 → 2026-07-31

Core

Provisioning, management, and uptime of CoreWeave's fleet of server nodes and supercomputing clusters.

Role type

Operations Engineer, Fleet Reliability

Builds

High-performance supercomputing clusters with state-of-the-art GPUs for AI workloads

Domain

Cloud Infrastructure / HPC / AI Compute

Deliverable

infrastructure

Required skills

Linux system administration, hardware troubleshooting, software troubleshooting, scripting (bash, python, powershell)

Preferred skills

Observability platforms (Grafana, Prometheus), Kubernetes administration, HPC/GPU workload administration, data center environment experience

Technologies

Linux, Kubernetes, Grafana, Prometheus, bash, python, powershell

Responsibilities

Configure and maintain large-scale high-performance supercomputing clusters; Troubleshoot hardware and software issues; Monitor and analyze system performance; Create and maintain documentation of team processes; Participate in oncall rotations

Seniority

Mid-level, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗