Site Reliability Engineer - Lead
Core
Lead Site Reliability Engineering for large-scale, distributed, fault-tolerant systems, ensuring reliability and performance of internal and external services.
Role type
Senior IC Site Reliability Engineer (Lead)
Builds
Cloud-native and hybrid infrastructure, CI/CD pipelines, automated tooling, and runbooks for production systems.
Domain
Financial services, cloud infrastructure, distributed systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Python, Bash, Java, Go, JavaScript, Node.js, Terraform, Chef, Ansible, Docker, Kubernetes, Linux/Windows administration, CI/CD, cloud architecture, incident management, postmortem analysis
Preferred skills
Cloud certification, DevSecOps practices, systems thinking, technical communication
Technologies
AWS, GCP, Jenkins, Terraform, Docker, Kubernetes, Python, Bash, Java, Go, JavaScript, Node.js
Responsibilities
Manage system uptime across cloud-native and hybrid architectures; Build CI/CD pipelines for application and cloud architecture patterns; Build automated tooling and comprehensive runbooks for service deployment and remediation; Solve problems and triage complex distributed architecture service maps; Lead availability blameless postmortems and drive remediation actions
Seniority
Senior, hands-on IC with leadership responsibilities