Staff Site Reliability Engineer
Core
Design, build, and evolve foundational systems, tooling, and operational practices to enable engineering teams to ship secure, reliable, and scalable software on AWS and Kubernetes.
Role type
Staff Site Reliability Engineer (Cloud-Native)
Builds
Scalable distributed systems, automated delivery pipelines, and observability tooling for mission-critical services
Domain
Fintech / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
AWS services (EKS, IAM, VPC, Lambda, CloudFront, S3), Kubernetes (operating and scaling), Istio service mesh, Infrastructure as Code (AWS CDK), CI/CD (GitHub Actions), TypeScript, Node.js, distributed systems troubleshooting, incident management, SLO/SLI definition, zero-trust architecture, mTLS, service networking, resilience engineering (autoscaling, failure testing)
Preferred skills
Progressive delivery practices (canary, blue/green), regulated/compliance-driven environment experience, FinOps principles, internal developer platform development, cloud-native certifications (CKA, CKAD, CKS, KCSA, KCNA)
Technologies
AWS, Kubernetes, Istio, AWS CDK, GitHub Actions, TypeScript, Node.js, OpenTelemetry, AWS X-Ray, CloudWatch
Responsibilities
Drive reliability engineering initiatives and operational excellence for mission-critical services; Design and improve deployment, release, and rollback strategies; Establish secure-by-default CI/CD pipelines; Enhance platform observability through metrics, logs, and tracing; Define and mature SLOs and reliability standards; Lead high-severity incident response and post-incident reviews; Partner with engineering teams to optimize runtime performance and embed reliability best practices; Mentor engineers on cloud-native technologies and SRE principles
Seniority
Staff, hands-on IC with mentorship