Lead Site Reliability Engineer (Kubernetes Required) - Hybrid
Core
Ensure reliability, scalability, and performance of production systems and services for investment professionals.
Role type
Lead Site Reliability Engineer (Kubernetes)
Builds
Robust infrastructure, automated processes, and reliable services for financial data platforms.
Domain
Financial technology / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes cluster management, Helm, CI/CD tooling, Infrastructure as Code, Monitoring & Observability, Python/Go/Bash scripting
Preferred skills
Open-source contributions, Google SRE principles, DevOps/Platform Engineering experience
Technologies
Kubernetes, AWS, GCP, Azure, GitHub Actions, ArgoCD, Harness, Prometheus, Grafana, Coralogix, OpenTelemetry, Terraform, Pulumi, Ansible, Puppet, Chef
Responsibilities
Monitor and improve production system reliability, respond to and resolve incidents, define and track SLOs/SLIs, design automation to reduce toil, participate in on-call rotation, contribute to capacity planning, document systems and runbooks
Seniority
Lead, hands-on IC with mentorship