Member of Technical Staff | Inference Platform
Core
Own the inference platform to execute models for enterprise and mid-market customers via large batch jobs and real-time APIs, ensuring fast, predictable, and cost-efficient serving.
Role type
Senior IC platform engineer (inference runtime)
Builds
Sophos (online and batch inference runtime on Kubernetes), controllers, and custom resources
Domain
Cloud infrastructure, Kubernetes, large-scale model serving, and data processing
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Kubernetes controllers/operators, Python, profiling and optimization, cost optimization, telemetry, batch scheduling, GPU serving
Preferred skills
Ray/Ray Serve/KubeRay, Kueue, Lance/Arrow/Parquet, GCP/AWS/GKE/EKS
Responsibilities
Evolve Sophos inference runtime, run large batch inference with multi-dimensional admission control, build and extend Sophos controller, optimize inference engines and feature processing, serve graphs and data from Lance-based storage, own execution of training and fine-tuning jobs, drive autoscaling and performance optimization, solve open problems in job sizing and resilience
Seniority
Senior, hands-on IC