CareerPlanGet AI match score →

Principal Engineer, AI Inference Reliability

US and Canada Offices💼 Full-time🗓 2025-10-29 → 2026-07-31

Core

Own the mission of making Cerebras Inference the most reliable AI service by driving reliability strategy and execution across the inference stack from client SDKs to wafer-scale systems.

Role type

Principal Reliability Tech Lead (IC)

Builds

Cerebras Inference service (ultra high-speed AI inference)

Domain

AI Infrastructure / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

SLO/SLI/SLA design, incident response, postmortem culture, fault detection, graceful degradation, failover, throttling, recovery, chaos testing, load simulation, distributed fault injection, system architecture for redundancy and durability

Preferred skills

building large-scale AI infrastructure systems

Technologies

Python, C++, Go, Rust

Responsibilities

Define and drive reliability strategy including SLOs; Design and implement reliability mechanisms for fault detection and recovery; Lead large-scale incident management and root-cause analysis; Architect for reliability and observability; Develop reliability tooling for chaos testing and fault injection; Collaborate across software, infrastructure, and hardware teams; Monitor and communicate reliability metrics; Mentor engineers on best practices for reliable system design

Seniority

Principal, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗