Lead Site Reliability Engineer
Core
Designing and building automated, self-service infrastructure platforms to manage scale, reliability, and security for a large cloud-native footprint.
Role type
Lead Site Reliability Engineer (Infrastructure)
Builds
Automated provisioning, upgrade, and migration systems for 15+ Kubernetes clusters and 100+ databases.
Domain
Fintech / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Infrastructure as Code (Terraform), Kubernetes, distributed systems debugging, observability tools (Datadog, CloudWatch, ELK), Python/Go/JavaScript, production code review, incident response strategy
Preferred skills
Large-scale AWS production Kubernetes, internal developer platform building, distributed systems architecture
Technologies
Kubernetes, Terraform, Docker, Datadog, CloudWatch, ELK, Python, Go, JavaScript, AWS
Responsibilities
Design systems to automate infrastructure management; reduce operational toil via repeatable workflows; build internal tooling for safe self-service changes; improve reliability and resilience of infrastructure; implement application deployment systems in Kubernetes; contribute to architecture decisions; write production-quality code
Seniority
Senior, hands-on IC with leadership scope