Sr Engineer Site Reliability
Core
Technical leader driving reliability initiatives, architecting solutions for complex operational challenges, and mentoring engineers to ensure platform scalability for millions of customers.
Role type
Senior Site Reliability Engineer (Technical Lead)
Builds
Highly available, fault-tolerant financial services infrastructure and automation frameworks
Domain
Financial Services / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
AWS architecture, Kubernetes (custom resources/operators), Terraform, Python/Go, Observability (Datadog/Splunk), CI/CD, GitOps, Incident response, Postmortems, Networking, Security, Mentoring
Preferred skills
Financial services domain knowledge, Compliance frameworks (SOC 2/PCI DSS/FINRA), AWS/DevOps certifications, Service mesh (Istio/Linkerd), Chaos engineering, FinOps
Technologies
AWS, EKS, Kubernetes, Terraform, Datadog, Splunk, ArgoCD, GitLab CI, Jenkins, Python, Go, Helm, Prometheus, Grafana, Istio
Responsibilities
Design and implement highly available, fault-tolerant systems; Architect infrastructure solutions using AWS best practices; Lead complex incident response efforts; Drive postmortem processes for high-severity incidents; Establish and track SLOs and SLIs; Build sophisticated Infrastructure as Code (IaC) solutions; Architect and optimize multi-cluster EKS environments; Design observability strategies; Implement progressive delivery mechanisms; Partner with development teams to improve application reliability; Mentor and guide junior and intermediate SREs; Implement and maintain zero-trust security controls
Seniority
Senior, hands-on IC with technical leadership