Site Reliability Engineer
Core
Design, implement, and maintain highly available, scalable, and reliable cloud infrastructure and services using automation and infrastructure-as-code.
Role type
Senior Site Reliability Engineer (IC)
Builds
Scalable cloud infrastructure, container-based applications, and observability solutions on AWS
Domain
Cloud Infrastructure / Platform Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes, Terraform, AWS, Go or Python, observability platforms (OpenTelemetry, Prometheus, Grafana, ELK/Splunk), incident response, SLOs
Preferred skills
Distributed systems, microservices, service mesh, high compliance environments, OpenTelemetry instrumentation, financial technology experience
Technologies
AWS, Kubernetes, Terraform, Spacelift, Pulumi, OpenTelemetry, Prometheus, Grafana, Jaeger, ELK, Splunk
Responsibilities
Design and implement comprehensive observability solutions; Build and maintain scalable cloud infrastructure; Automate routine operational tasks; Conduct post-incident reviews and root cause analysis; Participate in 24/7 on-call rotation; Scale container-based applications
Seniority
Senior, hands-on IC