CareerPlanGet AI match score →

Staff Site Reliability Engineer (Production Engineer)

Office - San Jose, USA💼 Full-time💰 $119,000–$119,000🗓 2026-07-14 → 2026-07-31

Core

Own the reliability of a large-scale cloud service (Linux/BSD, bare metal, Kubernetes, custom load balancing, SD-WAN) ensuring availability, latency, performance, efficiency, and scalability for a cloud processing tens of billions of transactions daily.

Role type

Staff Site Reliability Engineer (Production Engineer)

Builds

Cloud-native Zero Trust Exchange platform

Domain

Cybersecurity / Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Linux/Unix systems fundamentals, networking protocols (HTTP, DNS, TCP/IP, ICMP, OSI), programming (Python, Bash, Go), incident response, troubleshooting, CI/CD, capacity tuning, OS/app upgrades, vulnerability patching

Preferred skills

AI/ML frameworks or AIOps tools, Kubernetes at scale, Prometheus/OpenTelemetry ecosystems

Technologies

Linux, BSD, Kubernetes, SD-WAN, Prometheus, OpenTelemetry, Python, Bash, Go

Responsibilities

Define requirements for platform resilience, develop and operate end-to-end observability, lead full-cycle incident response, build and maintain everything-as-code for fleet lifecycle, continuously improve platform hygiene

Seniority

Staff, hands-on IC with strategic impact

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗