Director, Site Reliability Engineering
Core
Lead SRE teams and technical strategy to ensure critical infrastructure reliability, scalability, and security for millions of users.
Role type
Director, Site Reliability Engineering
Builds
Large-scale, privacy-focused technology environment
Domain
Cloud infrastructure, distributed systems, automation
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, platform engineering, Linux administration, distributed systems, cloud-native architectures, infrastructure automation, root-cause analysis, high-level programming (Go, Python, TypeScript), AI-driven software development, Docker
Preferred skills
Strategic thinking, technical foresight, complex project leadership
Technologies
Docker, Docker Compose, Go, Perl, TypeScript, Python
Responsibilities
Lead and develop SRE teams for large-scale system health; Define technical direction for infrastructure and deployment; Investigate and resolve instability in high-traffic distributed systems; Establish monitoring, alerting, and incident-response processes; Drive automation for infrastructure provisioning; Partner with software engineers on production issues; Guide long-term evolution of infrastructure architecture.
Seniority
Director, strategic leadership with hands-on technical depth