CareerPlanSign in

Machine Learning Performance Engineer - Offboard Training & Inference

Sunnyvale💼 Full-time🗓 2026-08-12 → 2026-09-26

Core

Optimizing large-scale distributed machine learning training and high-throughput batch inference workloads for petabyte-scale autonomy logs to maximize throughput, cluster goodput, and cost efficiency.

Role type

Senior Machine Learning Performance Engineer (Distributed Systems)

Builds

Optimized training pipelines and offline inference systems for physical AI autonomy stacks

Domain

Physical AI / Autonomous Vehicles / High-Performance Computing

Deliverable

production ML models

Required skills

Distributed multi-node training optimization, GPU/accelerator performance tuning, High-throughput batch inference systems, Systems programming (C++/Python), Profiling and root-cause analysis, Cluster goodput optimization

Preferred skills

GPU kernel development (CUDA/Triton), Profiling toolchains (Nsight/PyTorch Profiler), GPU scheduling on Kubernetes/Slurm, Fault tolerance for long-running jobs

Technologies

PyTorch, DeepSpeed, FSDP, Megatron, NCCL, NVIDIA Triton, TensorRT, ONNX Runtime, Ray, CUDA, Kubernetes, Slurm

Responsibilities

Profile and optimize distributed training end-to-end including data loading, kernel execution, and gradient communication; Optimize large-scale offline inference over petabyte-scale sensor logs; Establish roofline models and quantify performance gaps; Improve multi-node scaling efficiency via sharding and collective communication; Drive cluster goodput by reducing GPU idle time and I/O stalls; Build benchmarking and observability tooling for performance regression detection.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.