CareerPlanGet AI match score →

Staff ML Engineer, Inference Platform

Sunnyvale, California, United States of America💼 Full-time💰 $185,500–$185,500🗓 2026-05-28 → 2026-08-01

Core

Design and scale robust cloud-agnostic compute platforms for serving state-of-the-art machine learning models in real-time and batch scenarios for autonomous vehicles and AI-driven products.

Role type

Staff ML Infrastructure Engineer (Inference Platform)

Builds

Cloud-agnostic, reliable, and cost-efficient ML inference platform supporting experimental and bulk inference.

Domain

Automotive AI / Machine Learning Infrastructure

Deliverable

production ML models

Required skills

Distributed systems design, ML inference and model serving frameworks (Triton, RayServe, vLLM), High-performance backend development (Go, Python, C++), Cloud platforms (GCP, Azure, AWS), GPU utilization optimization, System observability and metrics.

Preferred skills

Building ML infrastructure platforms, Designing APIs and clients for ML workflows, Ray framework, Large-scale data processing, Telemetry and feedback loops, Hardware acceleration optimizations, Open-source contributions.

Technologies

Triton, RayServe, vLLM, Ray, B200, H100, A100, GCP, Azure, AWS

Responsibilities

Design and implement core platform backend software components, Lead technical decision-making on model serving strategies and orchestration, Drive development of monitoring and observability solutions, Proactively research and integrate state-of-the-art serving frameworks, Lead large-scale technical initiatives across the ML ecosystem, Establish best practices and contribute to open source projects.

Seniority

Staff, hands-on IC with technical leadership

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗