Principal SRE - AI Inference
Core
Architecting self-service platforms and control planes for ultra-reliable, large-scale AI inference infrastructure on Wafer-Scale Engine chips.
Role type
Principal Site Reliability Engineer (AI Inference Infrastructure)
Builds
Unified capacity management and production control plane for frontier AI model builders
Domain
AI/ML Inference Infrastructure, Wafer-Scale Engine (WSE), Large-scale Compute Fleets
Deliverable
production ML models | infrastructure
Required skills
Technical architecture for large-scale compute fleets, Capacity orchestration, Self-service platform design, SLO/SLI definition and management, Chaos engineering, Incident response leadership, Cross-functional stakeholder influence, Toil reduction strategy
Preferred skills
Bazel build systems, AI/ML inference system background, Predictive autoscaling, Cost-aware capacity management
Technologies
Wafer-Scale Engine (WSE), Multi-datacenter environments, Cloud-based solutions
Responsibilities
Define and implement robust strategies for delivering software reliably at scale across multiple datacenters, Architect self-service platforms for critical workflows, Define and evolve reliability practices including SLOs, SLIs, error budgets, and blameless postmortems, Mentor senior SREs and support critical incident escalations, Measure and drive impact through metrics like toil reduction, deployment velocity, and SLO compliance
Seniority
Principal, hands-on IC with strategic leadership