Site Reliability Engineer (SRE/ DevOps) - Engineering Productivity - Sydney
Core
Design, build, and operate secure, scalable, and fault-tolerant tools and infrastructure in a hybrid cloud environment to support Arista's product development teams.
Role type
Site Reliability Engineer (SRE) / DevOps Engineer
Builds
Internal CI/CD, testing, analysis, and visualization systems; production infrastructure and developer experience tools.
Domain
Networking hardware manufacturing; Cloud infrastructure; DevOps tooling
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Go, Python, shell scripting, Linux administration, server provisioning, infrastructure-as-code, problem solving, software troubleshooting
Preferred skills
Database management (MariaDB, Postgres, MongoDB), containerization (Docker, KVM, Kata-containers), monitoring stacks (Prometheus, Loki, Grafana), ElasticSearch, Artifactory, CI/CD systems (ArgoCD, Spinnaker), version control (Perforce, Gerrit), storage infrastructure (NAS, SAN, Ceph)
Technologies
Ansible, Artifactory, Gerrit, Jenkins, Kubernetes, Grafana, Spinnaker, MySQL, ElasticSearch, Google Cloud, Varnish, Perforce, Prometheus, Loki, Tempo, InfluxDB, Thanos, ArgoCD, Ceph
Responsibilities
Build, deploy, and operate critical production systems with focus on scalability, reliability, and observability; Monitor, support, and enhance developer experience across services; Build automation to remove toil; Create and maintain incident response runbooks; Triage platform/infrastructural issues; Engage with 3rd party vendor support; Plan and communicate maintenance windows; Design and implement solutions to resolve workflow bottlenecks.
Seniority
Mid-level, hands-on IC