CareerPlanGet AI match score →

Staff ML Performance Engineer (Training Efficiency)

Sunnyvale, California USA💼 Full-time💰 $336,400–$336,400🗓 2026-06-03 → 2026-07-31

Core

Optimizing large-scale ML training and inference workloads to enable scaling models to the next order of magnitude for embodied AI and automated driving systems.

Role type

Staff ML Performance Engineer (Training Efficiency)

Builds

Optimized training/inference pipelines for large-scale ML models

Domain

Embodied AI, Automated Driving, GPU Compute Infrastructure

Deliverable

production ML models

Required skills

Profiling ML workloads, designing efficiency improvements (parallelism, mixed precision), building observability tools, writing benchmarking tools, collaborating with research teams, high-quality Python development

Preferred skills

Concurrent/parallel/distributed computing, NVIDIA NSight Systems, GPU kernel implementation (CUDA, Triton), computing fundamentals

Technologies

NVIDIA Nsight Systems, CUDA, Triton, Python

Responsibilities

Profile ML workloads to identify bottlenecks; Design and implement efficiency improvements to maximize MFU and throughput; Design and implement observability tools to track performance metrics; Design and implement benchmarking tools; Collaborate with Research teams to integrate training efficiency improvements

Seniority

Staff, hands-on IC with strategic impact

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗