CareerPlanSign in

Cluster Operations Software Engineer

Headquarters/Sunnyvale Office💼 Full-time🗓 2026-08-13 → 2026-09-26

Core

Manage and operate cutting-edge machine learning compute clusters to ensure health, performance, and availability of Cerebras' Wafer-Scale Engine (WSE) infrastructure.

Role type

Senior IC cluster operations software engineer

Builds

Monitoring platforms, workflow automation systems, operational dashboards, reliability tooling, APIs, and integrations for global AI infrastructure

Domain

AI compute infrastructure / High-performance computing

Deliverable

production ML models | infrastructure

Required skills

Python, Go, Linux systems, Docker, Kubernetes, distributed systems, monitoring and alerting, incident response, resource optimization

Preferred skills

Large-scale AI cluster management, Ethernet/RoCE/TCP/IP, cloud platforms (AWS/GCP/Azure)

Technologies

Docker, Kubernetes, Python, Go, Linux, Ethernet, RoCE, TCP/IP, AWS, GCP, Azure

Responsibilities

Deploy and debug container-based services; build software solutions for cluster operations; develop APIs and automation services; manage and operate AI compute clusters; monitor cluster health and resolve issues; maximize compute capacity; handle engineering escalations

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.