Machine Learning Performance Engineer - Offboard Training & Inference
Core
Optimizing large-scale distributed machine learning training and high-throughput batch inference workloads for petabyte-scale autonomy logs to maximize throughput, cluster goodput, and cost efficiency.
Role type
Senior Machine Learning Performance Engineer (Distributed Systems)
Builds
Optimized training pipelines and offline inference systems for physical AI autonomy stacks
Domain
Physical AI / Autonomous Vehicles / High-Performance Computing
Deliverable
production ML models
Required skills
Distributed multi-node training optimization, GPU/accelerator performance tuning, High-throughput batch inference systems, Systems programming (C++/Python), Profiling and root-cause analysis, Cluster goodput optimization
Preferred skills
GPU kernel development (CUDA/Triton), Profiling toolchains (Nsight/PyTorch Profiler), GPU scheduling on Kubernetes/Slurm, Fault tolerance for long-running jobs
Technologies
PyTorch, DeepSpeed, FSDP, Megatron, NCCL, NVIDIA Triton, TensorRT, ONNX Runtime, Ray, CUDA, Kubernetes, Slurm
Responsibilities
Profile and optimize distributed training end-to-end including data loading, kernel execution, and gradient communication; Optimize large-scale offline inference over petabyte-scale sensor logs; Establish roofline models and quantify performance gaps; Improve multi-node scaling efficiency via sharding and collective communication; Drive cluster goodput by reducing GPU idle time and I/O stalls; Build benchmarking and observability tooling for performance regression detection.
Seniority
Senior, hands-on IC