Machine Learning Engineer, Performance Tooling
Core
Building and extending an end-to-end ML compilation pipeline to deploy models onto embedded hardware (NVIDIA, Qualcomm) for autonomous driving.
Role type
Staff-level ML Compiler Engineer
Builds
Deployable model bundles for NVIDIA TensorRT and Qualcomm QNN targets
Domain
Autonomous driving, embedded systems, compiler infrastructure
Deliverable
production ML models
Required skills
ML compilation pipeline ownership, graph lowering, quantization (PTQ), precision typing, graph partitioning, Python, C++
Preferred skills
MLIR, ONNX, TensorRT, Qualcomm QNN, PyTorch graph capture/export
Technologies
MLIR, ONNX, TensorRT, Qualcomm QNN, PyTorch
Responsibilities
Design and implement compiler passes with accuracy and latency gates, build scalable compilation infrastructure across platforms, partner with model teams on compilability, build regression and benchmarking suites, set technical direction and mentor on compiler design
Seniority
Staff, greenfield ownership