Staff ML Performance Engineer (Inference Optimisation)
Core
Optimizing ML inference for edge accelerators and GPUs to run large transformer-based models efficiently on low-cost, low-power edge devices for in-vehicle compute.
Role type
Staff ML Performance Engineer (Inference Optimisation)
Builds
Production ML systems running on edge devices (NVIDIA Orin/Thor, Qualcomm)
Domain
Automotive / Edge AI / Embedded Systems
Deliverable
production ML models
Required skills
Proficiency with ML inference stacks (TensorRT, CUDA, Triton, OpenCL), low-level kernel/runtime optimization, profiling and bottleneck analysis, benchmarking and regression testing, C++ and Python
Preferred skills
Embedded/edge deployment experience, NVIDIA/Qualcomm SoC expertise, mentoring, technical roadmap planning
Technologies
TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL, NVIDIA Orin, NVIDIA Thor, Qualcomm SoCs
Responsibilities
Profile and pinpoint bottlenecks across the full inference stack; Implement and validate optimisations in compilers, runtimes, and kernels; Build robust benchmarking and regression testing; Optimise for multiple hardware targets; Collaborate with model developers to influence architecture and deployment decisions; Contribute to technical roadmaps and tooling
Seniority
Staff, hands-on IC with technical direction