Staff ML Engineer, Inference Platform
Core
Design and scale robust cloud-agnostic compute platforms for serving state-of-the-art machine learning models in real-time and batch scenarios for autonomous vehicles and AI-driven products.
Role type
Staff ML Infrastructure Engineer (Inference Platform)
Builds
Cloud-agnostic, reliable, and cost-efficient ML inference platform supporting experimental and bulk inference.
Domain
Automotive AI / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
Distributed systems design, ML inference and model serving frameworks (Triton, RayServe, vLLM), High-performance backend development (Go, Python, C++), Cloud platforms (GCP, Azure, AWS), GPU utilization optimization, System observability and metrics.
Preferred skills
Building ML infrastructure platforms, Designing APIs and clients for ML workflows, Ray framework, Large-scale data processing, Telemetry and feedback loops, Hardware acceleration optimizations, Open-source contributions.
Technologies
Triton, RayServe, vLLM, Ray, B200, H100, A100, GCP, Azure, AWS
Responsibilities
Design and implement core platform backend software components, Lead technical decision-making on model serving strategies and orchestration, Drive development of monitoring and observability solutions, Proactively research and integrate state-of-the-art serving frameworks, Lead large-scale technical initiatives across the ML ecosystem, Establish best practices and contribute to open source projects.
Seniority
Staff, hands-on IC with technical leadership