Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
Core
Design, build, and scale control plane and data plane services for a distributed AI/ML cloud platform, ensuring high-performance execution of workloads across cloud and on-prem environments.
Role type
Senior Site Reliability Engineer (Platform Infrastructure)
Builds
Scalable, secure, and robust infrastructure for distributed AI/ML applications using Ray, Kubernetes, and cloud-native technologies.
Domain
Cloud-native distributed systems, Machine Learning infrastructure, Container orchestration
Deliverable
infrastructure
Required skills
Go, Python, Kubernetes, Cloud-native technologies (AWS, Azure, GCP), Distributed systems architecture, Networking, Security and authentication, Observability stacks (Prometheus, Grafana), Linux kernel and file systems, Container image management
Preferred skills
Experience with accelerator integration (GPUs, TPUs), Open-source contribution
Technologies
Ray, Kubernetes, AWS, Azure, GCP, Prometheus, Grafana, Go, Python
Responsibilities
Design and build services to orchestrate Ray clusters across cloud and on-prem environments; Optimize control plane components for large-scale distributed AI/ML workloads; Build intelligent scheduling and resource management systems; Develop features to enhance reliability, performance, scalability, and observability; Support and optimize accelerator integration; Handle container image management and dependency resolution; Provide on-call support for infrastructure issues.
Seniority
Senior, hands-on IC