Distinguished Site Reliability Engineer - Cloud
Core
Design, build, and maintain large-scale production GPU cloud services ensuring high efficiency, availability, and uptime through automation and systems engineering.
Role type
Distinguished Site Reliability Engineer (Cloud)
Builds
Large-scale Kubernetes clusters, GPU cloud services, and internal/external facing production systems
Domain
Cloud Infrastructure, Distributed Systems, GPU Computing
Deliverable
production ML models | infrastructure
Required skills
Linux, Networking, Containers, Python, Go, Perl, Ruby, Infrastructure automation, Distributed systems design, Capacity management, System design consulting, Real-time monitoring, Logging, Alerting, Incident response, Blameless postmortems
Preferred skills
Experience with OpenStack, Docker, Large-scale private/public cloud systems, Debugging and optimizing code, Automating routine tasks
Technologies
Kubernetes, OpenStack, Docker, Python, Go, Perl, Ruby
Responsibilities
Lead, design, implement, and support operational and reliability aspects of large-scale Kubernetes clusters; Engage in and improve the whole lifecycle of services from inception to refinement; Support services pre-launch through system design consulting and capacity management; Maintain live services by measuring availability, latency, and system health; Scale systems sustainably via automation; Practice sustainable incident response and blameless postmortems; Participate in on-call rotation for production systems
Seniority
Distinguished, hands-on IC with strategic impact