Team Leader, SRE
Core
Lead a Site Reliability Engineering team to ensure platform stability, manage incident response, and drive reliability practices for a global HR platform.
Role type
Senior IC Team Leader (SRE)
Builds
Kubernetes, AWS, PostgreSQL, CI infrastructure, and observability stack
Domain
Cloud Infrastructure / SRE / HR Tech
Deliverable
production ML models | infrastructure
Required skills
People leadership, Kubernetes production operations, AWS at scale, AI infrastructure enablement, Infrastructure as Code (Terraform), CI/CD systems, Docker, shell scripting, incident response, SLOs and error budgets, regulated environment experience
Preferred skills
Backend languages (Elixir, Java, Clojure, Node.js, Python), modern observability (OpenTelemetry, distributed tracing), PostgreSQL performance tuning, Linux systems outside cloud, security (defensive/offensive), FinOps, growing teams from scratch
Technologies
Kubernetes, AWS, PostgreSQL, Terraform, GitLab CI, GitHub Actions, Jenkins, Docker, OpenTelemetry, Honeycomb
Responsibilities
Manage full career lifecycle of direct reports (onboarding, feedback, performance, hiring), act as team spokesperson to engineering and leadership, define SRE goals and support rotation models, own core infrastructure reliability (K8s, AWS, Postgres, DNS, TLS), partner with Security team on threats and compliance, manage vendor relationships
Seniority
Senior, hands-on IC with leadership