Senior Site Reliability Engineer
Core
Build and maintain scalable infrastructure platforms supporting domestic and international workloads for a primary care AI tool, focusing on automation, container orchestration, and system reliability.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
Declarative application and infrastructure lifecycle management, continuous deployment pipelines, Kubernetes clusters, and automated delivery processes.
Domain
Healthcare technology / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Shell Scripting, Kubernetes, Docker, Containerd, Helm, gRPC, Prometheus, Linux system administration, TCP/IP, DNS, load balancing, public cloud platforms (GCP, Azure, AWS)
Preferred skills
CNCF-based technologies, networking fundamentals, SRE concepts (monitoring, performance tuning)
Responsibilities
Build systems for declarative application and infrastructure lifecycle management; prioritize and troubleshoot infrastructure issues to minimize downtime; streamline and automate infrastructure processes including delivery pipelines; collaborate with technical leads and data scientists to maintain scalable infrastructure.
Seniority
Senior, hands-on IC
