Lead Site Reliability Engineer (Kubernetes Required) - Hybrid
Core
Ensure reliability, scalability, and performance of production systems for financial data and analytics platforms serving investment professionals.
Role type
Lead Site Reliability Engineer (Kubernetes)
Builds
Robust infrastructure, automated processes, and reliable services for financial data platforms
Domain
Financial technology / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Kubernetes administration, incident response, SLO/SLI definition, automation design, capacity planning, system documentation
Preferred skills
Open-source contributions, Google SRE principles, DevOps/Platform Engineering experience
Technologies
Kubernetes, Helm, AWS/GCP/Azure, GitHub Actions/ArgoCD/Harness, Prometheus/Grafana/Coralogix/OpenTelemetry, Terraform/Pulumi, Ansible/Puppet/Chef, Python/Go/Bash
Responsibilities
Monitor and improve production system reliability, resolve incidents and conduct post-mortems, define and track SLOs/SLIs, collaborate on building reliability into services, design automation to reduce toil, participate in on-call rotation, contribute to capacity planning, document systems and runbooks
Seniority
Lead, hands-on IC with mentorship responsibilities