Software Engineer, Platform Infrastructure (Foundations)
Core
Building scalable, secure, and robust control and data plane infrastructure for distributed AI/ML workloads on the Ray platform.
Role type
Senior IC platform infrastructure engineer (distributed systems)
Builds
Scalable cluster orchestration services, intelligent scheduling systems, and high-performance execution environments for distributed AI/ML workloads.
Domain
Cloud-native distributed systems and machine learning infrastructure
Deliverable
production ML models | infrastructure
Required skills
Go, Python, Kubernetes, cloud-native technologies (AWS/Azure/GCP), distributed systems architecture, networking, security, Linux kernel, container orchestration, observability stacks (Prometheus/Grafana)
Preferred skills
GPU/TPU accelerator integration, heterogeneous compute cluster management, open-source contribution
Technologies
Ray, Kubernetes, Prometheus, Grafana, Linux, Go, Python
Responsibilities
Design and build services to orchestrate Ray clusters across cloud and on-prem environments; Optimize control plane components for large-scale distributed AI/ML workloads; Develop intelligent scheduling and resource management systems; Enhance reliability, performance, scalability, and observability of managed Ray workloads; Support and optimize accelerator integration (GPUs, TPUs); Handle container image management and dependency resolution; Provide on-call support for infrastructure issues.
Seniority
Senior, hands-on IC