Senior Software Engineer, Site Reliability Engineering
Core
Building resilient, highly available cloud platforms and automating operational challenges for mission-critical SaaS systems.
Role type
Senior Site Reliability Engineer
Builds
Cloud platforms, automation tooling, observability systems, and CI/CD pipelines for investment management SaaS.
Domain
Cloud Infrastructure / Investment Management
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
AWS services (EC2, ECS/EKS, RDS, S3, IAM, CloudWatch), observability (monitoring, alerting, distributed tracing), CI/CD pipeline design, deployment strategies (blue/green, canary), Infrastructure as Code (Terraform), scripting (Python, Go, Bash), incident response, SLO/SLI management
Preferred skills
Kubernetes, Helm, chaos engineering, SLO/error budget programs, Kotlin, Node.js, TypeScript, distributed cloud-native applications
Technologies
AWS, Terraform, GitHub Actions, CircleCI, Buildkite, Kubernetes, Helm
Responsibilities
Improve reliability and performance of production SaaS platform; Build automation to increase engineering velocity; Own and improve production observability; Design and enhance CI/CD pipelines; Define and improve SLOs and error budget practices; Identify capacity constraints and reliability risks; Participate in on-call rotation and incident response; Lead blameless postmortems; Partner with engineers on infrastructure design; Develop Infrastructure as Code solutions.
Seniority
Senior, hands-on IC