CareerPlanGet AI match score →

Principal Engineer, Inference Cloud

Headquarters/Sunnyvale Office💼 Full-time🗓 2025-09-29 → 2026-07-31

Core

Architecting and operating the cloud layer for Cerebras' ultra-high-speed AI inference service, ensuring multi-region availability, low latency, and reliability for model labs and enterprises.

Role type

Principal Engineer, Inference Cloud Platform

Builds

Multi-region cloud infrastructure for AI inference service

Domain

AI Infrastructure / Distributed Systems

Deliverable

infrastructure

Required skills

distributed systems architecture, cloud infrastructure, backend/systems languages (Go/C++/Python), high-availability system design, latency optimization, observability practices, production code contribution

Preferred skills

ML inference infrastructure, model serving systems, GPU-accelerated workloads, TTFT/tail-latency reduction

Technologies

Go, C++, Python

Responsibilities

Define and prioritize critical platform technical problems; set long-term technical direction for multi-region topology and service evolution; architect active-active systems with rapid failover and graceful degradation; contribute production code and review designs; lead resolution of hard production issues and drive operational rigor; drive platform-wide decisions on reliability and deployment strategy; mentor engineers on technical decision-making

Seniority

Principal, hands-on IC with strategic scope

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗