Cluster Operations Software Engineer
Core
Manage and operate cutting-edge machine learning compute clusters to ensure health, performance, and availability of Cerebras' Wafer-Scale Engine (WSE) infrastructure.
Role type
Senior IC cluster operations software engineer
Builds
Monitoring platforms, workflow automation systems, operational dashboards, reliability tooling, APIs, and integrations for global AI infrastructure
Domain
AI compute infrastructure / High-performance computing
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Linux systems, Docker, Kubernetes, distributed systems, monitoring and alerting, incident response, resource optimization
Preferred skills
Large-scale AI cluster management, Ethernet/RoCE/TCP/IP, cloud platforms (AWS/GCP/Azure)
Technologies
Docker, Kubernetes, Python, Go, Linux, Ethernet, RoCE, TCP/IP, AWS, GCP, Azure
Responsibilities
Deploy and debug container-based services; build software solutions for cluster operations; develop APIs and automation services; manage and operate AI compute clusters; monitor cluster health and resolve issues; maximize compute capacity; handle engineering escalations
Seniority
Senior, hands-on IC