Staff ML Performance Engineer (Compiler)
Core
Optimizing ML inference for edge accelerators and GPUs to run large transformer-based models efficiently on low-cost, low-power edge devices for automated driving systems.
Role type
Staff ML Performance Engineer (Compiler)
Builds
Production ML systems running reliably on in-vehicle compute
Domain
Autonomous driving, Edge AI, Compiler optimization
Deliverable
production ML models
Required skills
ML compiler optimization, runtime optimization, kernel development, profiling and bottleneck analysis, benchmarking and regression testing, multi-platform deployment (NVIDIA/Qualcomm), C++, Python
Preferred skills
Compute graph scheduling, embedded ML deployment, NVIDIA/Qualcomm SoC tooling, mentoring, technical roadmap planning
Technologies
TensorRT, CUDA, QNN, Triton, OpenCL, MLIR, ONNX, NVIDIA Orin, NVIDIA Thor, Qualcomm
Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.