Site Reliability Engineer - SRE (Platform Software Team)
Core
Develop and sustain fundamental infrastructure for HW Labs and Manufacturing sites to support validation of high-speed digital designs and optimized manufacturing yields for Arista Network products.
Role type
Site Reliability Engineer (Platform Software Team)
Builds
Critical production systems, automation workflows, and deployment pipelines for lab and manufacturing operations.
Domain
Computer Networking / Hardware Manufacturing / Data Centers
Deliverable
production ML models | product features | infrastructure
Required skills
Go, Python, bash shell scripting, Linux administration, Kubernetes on bare-metal, Ansible, Docker, virtualization, infrastructure-as-code, air-gapped systems management, incident response, postmortem analysis, system scaling, fault-tolerance implementation, OSS system design.
Preferred skills
MySQL/RDBMS management, Prometheus/Grafana monitoring stack, Artifactory/Docker registry management, CI/CD systems (GitLab CI/CD, Jenkins, Zuul), Terraform.
Responsibilities
Build, deploy, and operate critical production systems with focus on scalability, reliability, and observability; monitor and enhance product deployment experience; build automation to remove toil; proactively monitor and respond to alerts; create and maintain incident response runbooks; triage platform/infrastructural issues; plan and communicate maintenance windows; implement solutions to scale systems and improve availability; study OSS systems for better triage and fix resolution.
Seniority
Mid-Senior, hands-on IC
