CareerPlanGet AI match score →

Distinguished Site Reliability Engineer - Cloud

6 Locations💼 Full-time💰 $320,000–$320,000🗓 2026-07-10 → 2026-07-31

Core

Design, build, and maintain large-scale production GPU cloud services ensuring high efficiency, availability, and uptime through automation and systems engineering.

Role type

Distinguished Site Reliability Engineer (Cloud)

Builds

Large-scale Kubernetes clusters, GPU cloud services, and internal/external facing production systems

Domain

Cloud Infrastructure, Distributed Systems, GPU Computing

Deliverable

production ML models | infrastructure

Required skills

Linux, Networking, Containers, Python, Go, Perl, Ruby, Infrastructure automation, Distributed systems design, Capacity management, System design consulting, Real-time monitoring, Logging, Alerting, Incident response, Blameless postmortems

Preferred skills

Experience with OpenStack, Docker, Large-scale private/public cloud systems, Debugging and optimizing code, Automating routine tasks

Technologies

Kubernetes, OpenStack, Docker, Python, Go, Perl, Ruby

Responsibilities

Lead, design, implement, and support operational and reliability aspects of large-scale Kubernetes clusters; Engage in and improve the whole lifecycle of services from inception to refinement; Support services pre-launch through system design consulting and capacity management; Maintain live services by measuring availability, latency, and system health; Scale systems sustainably via automation; Practice sustainable incident response and blameless postmortems; Participate in on-call rotation for production systems

Seniority

Distinguished, hands-on IC with strategic impact

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗