Site Reliability Engineer (SRE/ DevOps) - Engineering Productivity
Core
Design, build, and operate secure, scalable, and fault-tolerant tools and infrastructure in a hybrid cloud environment to support Arista's product development teams.
Role type
Site Reliability Engineer (SRE) / DevOps Engineer
Builds
Internal developer platforms, CI/CD pipelines, monitoring stacks, and automation tools for software engineering teams.
Domain
Cloud Infrastructure / DevOps / Software Engineering Productivity
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Go, Python, shell scripting, Linux administration, infrastructure-as-code, server provisioning, problem solving, software troubleshooting
Preferred skills
Database management (MariaDB, PostgreSQL, MongoDB), containerization (Docker, KVM, Kata-containers), monitoring stacks (Prometheus, Loki, Grafana), CI/CD systems (ArgoCD, Spinnaker), version control (Perforce, Gerrit), storage infrastructure (NAS, SAN, Ceph)
Technologies
Ansible, Artifactory, Gerrit, Jenkins, Kubernetes, Grafana, Spinnaker, MySQL, ElasticSearch, Google Cloud, Varnish, Perforce, Prometheus, Loki, Tempo, InfluxDB, Thanos, ArgoCD, Ceph
Responsibilities
Build and operate critical production systems with focus on scalability, reliability, and observability; Monitor and enhance developer experience across services; Create and maintain incident response runbooks; Triage platform issues and assist software engineers; Engage with 3rd party vendor support; Plan and communicate maintenance windows; Identify and resolve infrastructural bottlenecks in developer workflows.
Seniority
Mid-level, hands-on IC