Principal Engineer, AI Inference Reliability
Core
Own the mission of making Cerebras Inference the most reliable AI service by driving reliability strategy and execution across the inference stack from client SDKs to wafer-scale systems.
Role type
Principal Reliability Tech Lead (IC)
Builds
Cerebras Inference service (ultra high-speed AI inference)
Domain
AI Infrastructure / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
SLO/SLI/SLA design, incident response, postmortem culture, fault detection, graceful degradation, failover, throttling, recovery, chaos testing, load simulation, distributed fault injection, system architecture for redundancy and durability
Preferred skills
building large-scale AI infrastructure systems
Technologies
Python, C++, Go, Rust
Responsibilities
Define and drive reliability strategy including SLOs; Design and implement reliability mechanisms for fault detection and recovery; Lead large-scale incident management and root-cause analysis; Architect for reliability and observability; Develop reliability tooling for chaos testing and fault injection; Collaborate across software, infrastructure, and hardware teams; Monitor and communicate reliability metrics; Mentor engineers on best practices for reliable system design
Seniority
Principal, hands-on IC