Senior Site Reliability Engineer
Core
Design and build robust, scalable, and fault-tolerant infrastructure and services for a high-throughput healthcare platform.
Role type
Senior Site Reliability Engineer (IC)
Builds
AWS-based platform, self-healing systems, CI/CD pipelines, observability tooling
Domain
Healthcare / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS, Kubernetes, distributed system architecture, infrastructure as code (Terraform), observability (Grafana, OpenTelemetry, Prometheus), incident response, CI/CD pipeline design, chaos engineering, disaster recovery
Preferred skills
Java, Python, Go, GitOps best practices, mentorship
Technologies
AWS, Kubernetes, Terraform, Grafana, OpenTelemetry, Prometheus, Datadog, Git
Responsibilities
Design and implement scalable, fault-tolerant infrastructure and services on AWS and Kubernetes; Define and help drive adoption of SLIs, SLOs, and SLAs; Own and improve observability using Grafana, OpenTelemetry, and related tooling; Build and maintain infrastructure as code (Terraform) and contribute to GitOps best practices; Participate in incident response and on-call rotation; Support the growth of junior and mid-level SRE engineers through mentorship, code reviews, and knowledge sharing; Drive continuous improvements in CI/CD pipelines, service ownership, chaos engineering, disaster recovery, and secure deployments.
Seniority
Senior, hands-on IC