Senior Site Reliability Engineer
Core
Own the reliability, performance, and resilience of cloud infrastructure powering Garner's products and AI/ML workloads, ensuring production quality for healthcare outcomes.
Role type
Senior Site Reliability Engineer (Platform Engineering)
Builds
Cloud environments (AWS, Kubernetes), observability systems, and automated infrastructure-as-code deliverables for AI/ML workloads.
Domain
Healthcare technology / Cloud Infrastructure / AI/ML
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, AWS, Python, Go, SLO definition, incident response, observability, infrastructure automation, cloud cost optimization
Preferred skills
AI/ML workload support, regulated environment operations (HIPAA/SOC 2), AI tool fluency
Technologies
AWS, Kubernetes, Terraform, Istio, Python, Go, TypeScript, Postgres, NATS, Datadog, GitLab
Responsibilities
Define and uphold SLOs across critical services; lead incident response and root cause analysis; build and maintain monitoring and alerting systems; translate scaling requirements into automated infrastructure-as-code; automate operational toil using AI tools; enable engineering team deployment standards; ensure HIPAA and security compliance
Seniority
Senior, hands-on IC