Principal Site Reliability Engineer
Core
Define technical vision and architecture for a self-healing, AI-assisted reliability and infrastructure platform serving thousands of companies.
Role type
Principal Site Reliability Engineer (Strategy & Architecture)
Builds
Kubernetes infrastructure, observability frameworks, SLO programs, incident management systems, and AI-assisted operational agents.
Domain
Cloud Infrastructure & Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes internals and large-scale cluster management, SRE principles (SLIs/SLOs/Error Budgets), distributed systems failure mode analysis, cloud networking and security, incident management leadership, CI/CD pipeline safety, AI-assisted operational system design.
Preferred skills
Experience shipping AI tooling in production, org-wide observability framework design, progressive delivery strategies.
Technologies
Kubernetes, OpenTelemetry, Prometheus, Grafana, Datadog, Elasticsearch, AWS, GCP, Azure.
Responsibilities
Define long-term technical architecture for reliability platform, lead design of major platform initiatives, partner with leadership on infrastructure decisions, identify systemic risks and drive org-wide remediation, operationalize SLIs/SLOs across engineering teams, design AI-assisted operational systems, mentor Staff/Senior SREs.
Seniority
Principal, hands-on IC with strategic leadership
