Staff Site Reliability Engineer (Production Engineer)
Core
Own the reliability of a large-scale cloud service (Linux/BSD, bare metal, Kubernetes, custom load balancing, SD-WAN) ensuring availability, latency, performance, efficiency, and scalability for a cloud processing tens of billions of transactions daily.
Role type
Staff Site Reliability Engineer (Production Engineer)
Builds
Cloud-native Zero Trust Exchange platform
Domain
Cybersecurity / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux/Unix systems fundamentals, networking protocols (HTTP, DNS, TCP/IP, ICMP, OSI), programming (Python, Bash, Go), incident response, troubleshooting, CI/CD, capacity tuning, OS/app upgrades, vulnerability patching
Preferred skills
AI/ML frameworks or AIOps tools, Kubernetes at scale, Prometheus/OpenTelemetry ecosystems
Technologies
Linux, BSD, Kubernetes, SD-WAN, Prometheus, OpenTelemetry, Python, Bash, Go
Responsibilities
Define requirements for platform resilience, develop and operate end-to-end observability, lead full-cycle incident response, build and maintain everything-as-code for fleet lifecycle, continuously improve platform hygiene
Seniority
Staff, hands-on IC with strategic impact