Team Leader, SRE
Core
Lead a Site Reliability Engineering team responsible for the reliability of critical cloud infrastructure and developer platforms, balancing 60% hands-on technical contribution with 40% people leadership.
Role type
Senior IC Team Leader, SRE
Builds
Cloud infrastructure, developer platforms, and reliability practices
Domain
Cloud infrastructure, Site Reliability Engineering, AI infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, AWS, Terraform, CI/CD, observability, incident response, SLOs, error budgets, team leadership, hiring, conflict resolution, backend programming (Elixir/Java/Clojure/Node.js/Python), PostgreSQL/Aurora operations, Linux administration, infrastructure security, FinOps
Preferred skills
AI infrastructure scaling, modern observability (OpenTelemetry, distributed tracing, Honeycomb), growing engineering teams from scratch
Technologies
Kubernetes, AWS, PostgreSQL, Terraform, GitLab CI, GitHub Actions, Jenkins, Docker, OpenTelemetry, Honeycomb, Elixir, Java, Clojure, Node.js, Python
Responsibilities
Lead and develop an SRE team, owning the full career lifecycle of direct reports; Coach engineers on technical craft and interpersonal skills; Foster a healthy, collaborative team environment; Represent the SRE team across engineering and with senior leadership; Define and prioritize SRE objectives; Own the support rotation and on-call model; Provide technical leadership across core infrastructure; Drive the evolution of reliability practices; Partner with Security teams on infrastructure threats; Oversee infrastructure-related vendor relationships; Stay hands-on to review technical work and guide decisions during incidents; Help mature the organization's reliability practices; Balance short-term operational demands with longer-term platform investments
Seniority
Senior, hands-on IC with leadership responsibilities