Principal Engineer, Inference Cloud
Core
Architecting and operating the cloud layer for Cerebras' ultra-high-speed AI inference service, ensuring multi-region availability, low latency, and reliability for model labs and enterprises.
Role type
Principal Engineer, Inference Cloud Platform
Builds
Multi-region cloud infrastructure for AI inference service
Domain
AI Infrastructure / Distributed Systems
Deliverable
infrastructure
Required skills
distributed systems architecture, cloud infrastructure, backend/systems languages (Go/C++/Python), high-availability system design, latency optimization, observability practices, production code contribution
Preferred skills
ML inference infrastructure, model serving systems, GPU-accelerated workloads, TTFT/tail-latency reduction
Technologies
Go, C++, Python
Responsibilities
Define and prioritize critical platform technical problems; set long-term technical direction for multi-region topology and service evolution; architect active-active systems with rapid failover and graceful degradation; contribute production code and review designs; lead resolution of hard production issues and drive operational rigor; drive platform-wide decisions on reliability and deployment strategy; mentor engineers on technical decision-making
Seniority
Principal, hands-on IC with strategic scope