Staff ML Performance Engineer (Training Efficiency)
Core
Optimizing large-scale ML training and inference workloads to enable scaling models to the next order of magnitude.
Role type
Staff ML Performance Engineer (Training Efficiency)
Builds
Optimized training/inference pipelines for large-scale ML models
Domain
Machine Learning, GPU Compute Infrastructure
Deliverable
production ML models
Required skills
Profiling ML workloads, designing efficiency improvements (parallelism, model compilation, mixed precision), building observability tools, benchmarking, Python development, distributed platforms
Preferred skills
Concurrent/parallel/distributed computing, NVIDIA NSight Systems, GPU kernels (CUDA, Triton), computing fundamentals
Technologies
NVIDIA Nsight Systems, CUDA, Triton, Python
Responsibilities
Profile ML workloads to identify bottlenecks; Design and implement efficiency improvements to maximize MFU and throughput; Design and implement observability tools to track performance metrics; Design and implement benchmarking tools; Collaborate with Research teams to integrate performance improvements
Seniority
Staff, hands-on IC with mentorship