Staff ML Performance Engineer (Inference Optimisation)
Core
Optimizing ML inference for edge accelerators and GPUs to run large transformer-based models efficiently on low-cost, low-power edge devices for automated driving systems.
Role type
Staff ML Performance Engineer (Inference Optimisation)
Builds
Production ML systems for in-vehicle compute
Domain
Autonomous driving / Edge AI / Embedded Systems
Deliverable
production ML models
Required skills
Proficiency with inference stacks (TensorRT, CUDA, Triton, OpenCL), profiling and bottleneck analysis, compiler/runtime optimization, kernel development, benchmarking and regression testing, C++ and Python
Preferred skills
Embedded or edge deployment experience, NVIDIA/Qualcomm SoC expertise, mentoring, technical roadmap planning
Technologies
TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL, NVIDIA Orin, NVIDIA Thor, Qualcomm
Responsibilities
Profile and pinpoint bottlenecks across the full inference stack; Implement and validate optimisations in compilers, runtimes, and kernels; Build robust benchmarking and regression testing; Optimise for multiple hardware targets; Collaborate with model developers to influence architecture and deployment decisions; Contribute to technical roadmaps and tooling
Seniority
Staff, hands-on IC with technical direction