Site Reliability Engineer Specialist
Core
Design, build, and maintain scalable infrastructure and systems to ensure services are reliable, efficient, and resilient.
Role type
Site Reliability Engineer
Builds
Scalable infrastructure, automated operational tasks, CI/CD pipelines, and disaster recovery strategies.
Domain
Cloud infrastructure and distributed systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Python, Go, Java, AWS, Azure, GCP, Terraform, Ansible, Linux/Unix, networking fundamentals, Docker, Kubernetes, CI/CD pipeline management, root cause analysis, incident management, system optimization, distributed system debugging
Preferred skills
Infrastructure as code, containerization and orchestration, cross-functional collaboration, continuous learning
Technologies
Python, Go, Java, AWS, Azure, GCP, Terraform, Ansible, Docker, Kubernetes
Responsibilities
Design and maintain scalable infrastructure, automate operational tasks, implement application monitoring, develop disaster recovery strategies, collaborate on application architecture, optimize system metrics, manage CI/CD pipelines, perform root cause analysis, consult on reliability best practices, maintain documentation, handle on-call shifts
Seniority
Mid to Senior, hands-on IC