Member of Technical Staff | Inference Platform
Core
Build and operate Kubernetes-based infrastructure for reliable, scalable machine learning inference across cloud and customer environments, handling both large-scale batch workloads and real-time APIs.
Role type
Senior IC systems engineer (ML inference platform)
Builds
Production ML inference runtime, batch execution systems, and GPU-optimized serving infrastructure
Domain
Cloud infrastructure + Machine Learning
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (controllers, operators, custom resources), Python (production systems), distributed systems, GPU optimization, data-intensive pipeline profiling, autoscaling strategies, telemetry/monitoring, security (encryption, isolation), cost-aware engineering
Preferred skills
Ray/Ray Serve/KubeRay, Kueue, columnar data formats (Arrow, Parquet, Lance), GCP/AWS (GKE/EKS), financial services experience
Technologies
Kubernetes, Python, Ray, Kueue, Lance, Arrow, Parquet, GKE, EKS
Responsibilities
Evolve and operate online and batch inference runtime; Implement multi-dimensional admission control for batch jobs; Build Kubernetes controllers for model scheduling; Optimize inference engines and feature-processing pipelines; Develop mechanisms for serving graphs from Lance-based storage; Design autoscaling and GPU serving strategies; Implement telemetry for performance and cost optimization; Solve complex challenges in job sizing, recovery, and profiling; Design secure execution approaches for per-customer encryption and isolation.
Seniority
Senior, hands-on IC