Software Engineer, Ray Serve
Core
Building production-grade distributed serving infrastructure for high-performance machine learning applications, enabling seamless deployment of complex AI models at scale.
Role type
Senior IC distributed systems engineer (ML serving infrastructure)
Builds
High-throughput ML serving frameworks, asynchronous inference engines, intelligent model routing systems, zero-downtime update mechanisms, and multi-model orchestration pipelines.
Domain
Distributed systems, high-performance computing, machine learning infrastructure
Deliverable
production ML models
Required skills
Deep systems fundamentals (OS, networking, concurrency), distributed systems architecture, production system design at scale, code quality and testing, ownership of full lifecycle from design to incident response
Preferred skills
Experience with gRPC and Ray, ML/AI systems background, open source contributions, performance optimization and profiling, cloud-native technologies (Kubernetes, Istio)
Technologies
Python, Cython, C++, Ray Core, gRPC, Kubernetes, TensorFlow, PyTorch, JAX, OpenTelemetry, Prometheus
Responsibilities
Design and implement asynchronous inference systems, build sub-millisecond model routing logic, develop zero-downtime traffic management for model updates, architect state management for large-scale replica clusters, create multi-model orchestration frameworks, build observability and debugging tools for distributed applications
Seniority
Senior, hands-on IC