Site Reliability Engineer
Core
Ensuring the reliability, availability, performance, and scalability of cloud infrastructure and healthcare platforms supporting millions of users.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Reliable, scalable, and secure cloud infrastructure and operational platforms for healthcare products.
Domain
Healthcare technology / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering (SRE) practices, Cloud infrastructure administration, Infrastructure as Code (IaC), Monitoring and observability, Incident management, Capacity planning, Automation scripting, Kubernetes management, CI/CD pipelines, Networking and security principles
Preferred skills
Disaster recovery planning, Business continuity strategies, Performance tuning, Operational runbook development
Technologies
AWS, Azure, Google Cloud Platform, Terraform, Bicep, ARM, CloudFormation, Docker, Kubernetes, PowerShell, Bash, Python
Responsibilities
Design and maintain reliable, scalable, and secure infrastructure; Define and monitor SLIs, SLOs, and SLAs; Build and maintain observability solutions; Participate in incident response and root cause analysis; Lead initiatives to reduce operational toil through automation; Collaborate with engineering teams to improve application reliability; Conduct capacity forecasting and scalability planning; Develop operational runbooks and best practices; Contribute to disaster recovery and business continuity initiatives.
Seniority
Mid-Senior, hands-on IC