Perception Deployment Engineer - Model Deployment & Optimization
Core
Deploying and optimizing large-scale multi-modal foundation models (sensor fusion, LLMs, VLMs) for real-time execution on vehicle SOCs.
Role type
Senior IC perception deployment engineer (model optimization & edge deployment)
Builds
Production-ready, low-latency inference code and optimized model binaries for autonomous vehicle stacks
Domain
Autonomous driving, computer vision, edge AI
Deliverable
production ML models
Required skills
C++ (14/17/20), CUDA, model quantization (PTQ, QAT), TensorRT, mixed-precision inference, custom ML OPs, PyTorch, ONNX, latency benchmarking
Preferred skills
FlashAttention, KV-cache optimization, BEV, 3D Occupancy Networks, VLM/VLA models, TensorRT-LLM
Responsibilities
Design and develop production-level C++ and CUDA code for real-time perception algorithms; Optimize large-scale models using quantization and mixed-precision frameworks; Architect and implement model conversion and compilation pipelines using TensorRT; Perform parity checking, accuracy recovery, and latency benchmarking; Develop and optimize custom ML OPs and TensorRT Plugins with efficient CUDA kernels