CareerPlanGet AI match score →

Principal SRE - AI Inference

Headquarters/Sunnyvale Office💼 Full-time🗓 2026-07-06 → 2026-07-31

Core

Architecting self-service platforms and control planes for ultra-reliable, large-scale AI inference infrastructure on Wafer-Scale Engine chips.

Role type

Principal Site Reliability Engineer (AI Inference Infrastructure)

Builds

Unified capacity management and production control plane for frontier AI model builders

Domain

AI/ML Inference Infrastructure, Wafer-Scale Engine (WSE), Large-scale Compute Fleets

Deliverable

production ML models | infrastructure

Required skills

Technical architecture for large-scale compute fleets, Capacity orchestration, Self-service platform design, SLO/SLI definition and management, Chaos engineering, Incident response leadership, Cross-functional stakeholder influence, Toil reduction strategy

Preferred skills

Bazel build systems, AI/ML inference system background, Predictive autoscaling, Cost-aware capacity management

Technologies

Wafer-Scale Engine (WSE), Multi-datacenter environments, Cloud-based solutions

Responsibilities

Define and implement robust strategies for delivering software reliably at scale across multiple datacenters, Architect self-service platforms for critical workflows, Define and evolve reliability practices including SLOs, SLIs, error budgets, and blameless postmortems, Mentor senior SREs and support critical incident escalations, Measure and drive impact through metrics like toil reduction, deployment velocity, and SLO compliance

Seniority

Principal, hands-on IC with strategic leadership

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗