CareerPlanGet AI match score →

Senior Site Reliability Engineer, Platform Infrastructure (Foundations)

San Francisco💼 Full-time🗓 2026-06-17 → 2026-07-31

Core

Design, build, and scale control plane and data plane services for a distributed AI/ML cloud platform, ensuring high-performance execution of workloads across cloud and on-prem environments.

Role type

Senior Site Reliability Engineer (Platform Infrastructure)

Builds

Scalable, secure, and robust infrastructure for distributed AI/ML applications using Ray, Kubernetes, and cloud-native technologies.

Domain

Cloud-native distributed systems, Machine Learning infrastructure, Container orchestration

Deliverable

infrastructure

Required skills

Go, Python, Kubernetes, Cloud-native technologies (AWS, Azure, GCP), Distributed systems architecture, Networking, Security and authentication, Observability stacks (Prometheus, Grafana), Linux kernel and file systems, Container image management

Preferred skills

Experience with accelerator integration (GPUs, TPUs), Open-source contribution

Technologies

Ray, Kubernetes, AWS, Azure, GCP, Prometheus, Grafana, Go, Python

Responsibilities

Design and build services to orchestrate Ray clusters across cloud and on-prem environments; Optimize control plane components for large-scale distributed AI/ML workloads; Build intelligent scheduling and resource management systems; Develop features to enhance reliability, performance, scalability, and observability; Support and optimize accelerator integration; Handle container image management and dependency resolution; Provide on-call support for infrastructure issues.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗