Senior Site Reliability Engineer - CloudVision
Core
Design, build, and operate scalable, reliable, and secure production systems and automation solutions for cloud networking infrastructure.
Role type
Senior Site Reliability Engineer (Cloud/Infrastructure)
Builds
Production systems, automation workflows, and monitoring infrastructure for cloud networking platforms.
Domain
Cloud networking, data center infrastructure, and distributed systems.
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Go, Python, bash shell scripting, Linux/UNIX administration, infrastructure-as-code, incident response, postmortem analysis, server provisioning, distributed systems architecture, performance tuning, security best practices.
Preferred skills
Kubernetes, Docker, virtualization, Prometheus, Grafana, CI/CD (GitLab, Spinnaker), Terraform, PostgreSQL, cloud platforms (AWS, GCP, Azure), artifact repositories, Docker registries.
Technologies
Ansible, Terraform, Prometheus, Grafana, AWS, GCP, Azure, Kubernetes, Docker, Jenkins, GitLab, Spinnaker, PostgreSQL.
Responsibilities
Design and deploy production systems with focus on scalability and security; develop automation solutions to reduce toil; monitor systems and implement automated incident response; create incident runbooks and conduct postmortems; collaborate with engineering teams to resolve bottlenecks; optimize monitoring infrastructure; execute maintenance windows; triage platform issues; deploy updates in staged, risk-managed rollouts.
Seniority
Senior, hands-on IC