Staff DevOps Engineer
Core
Ensuring reliability, scalability, and operational excellence for Zoom's real-time communications platform supporting audio/video conferencing, recording, and live-streaming.
Role type
Staff Site Reliability Engineer (SRE)
Builds
Global, large-scale distributed real-time media systems
Domain
Telecommunications / Real-time media infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE principles, SLO/SLI framework definition, incident response leadership, chaos engineering, observability tooling, infrastructure-as-code, CI/CD pipeline design, capacity planning, network architecture, cross-functional technical leadership, real-time communication protocols (WebRTC, RTP/RTCP, TURN/STUN), cloud infrastructure, container orchestration, Python/Bash/Go scripting
Preferred skills
Experience with media-heavy platforms (video conferencing, live streaming, gaming), architectural authority on deployment patterns, global team collaboration, mentoring senior engineers
Technologies
Kubernetes, Helm, ArgoCD, Terraform, Pulumi, Prometheus, Grafana, Datadog, Jaeger, OpenTelemetry, GitHub Actions, Jenkins, Spinnaker, AWS, GCP, Azure, WebRTC, BGP, DNS, CDN
Responsibilities
Own the SLO/SLI framework for real-time services; lead incident response for critical outages; implement chaos engineering and game day exercises; build and evolve observability tools; serve as architectural authority on deployment patterns and infrastructure design; drive capacity planning and cost optimization; establish DevOps best practices (IaC, GitOps); guide senior engineers on SRE principles; act as technical liaison between US and China/India teams; conduct architecture reviews and planning sessions in English and Mandarin.
Seniority
Staff, hands-on IC with strategic scope