Distributed Software Engineer
Core
Building and operating production distributed systems and infrastructure software for clusters of Cerebras AI chips, servers, and switches.
Role type
Senior IC distributed systems engineer (infrastructure)
Builds
Cloud-scale clusters of wafer-scale AI systems, Kubernetes operators, and control-plane services
Domain
AI hardware infrastructure, distributed systems, cloud operations
Deliverable
production ML models | infrastructure
Required skills
Go, Python, Kubernetes (operators, CRDs, reconciliation), distributed systems debugging, Linux, networking, Prometheus, Grafana
Preferred skills
bare-metal/HPC fleet operations, scheduler internals, RDMA/RoCE, eBPF, Ceph, NVMe-oF, etcd
Technologies
Go, Python, Kubernetes, gRPC, Prometheus, Grafana, Redfish, IPMI, gNMI, sFlow
Responsibilities
Automate bare-metal networking, OS, and application software across clusters; implement Kubernetes operators for large inference workloads; build gRPC control-plane services with authorization and quota policy; design metrics and log pipelines with custom exporters; implement failure detection, HA control planes, and automated recovery; develop CLIs, APIs, and MCP gateways for fleet access.
Seniority
Senior, hands-on IC