Senior AI Inference Engineer - Model Optimization & Deployment
Core
Optimizing and deploying large-scale multi-modal foundation models (LLMs, VLMs, sensor fusion) for real-time inference on power- and thermal-constrained vehicle SOCs.
Role type
Senior IC model optimization and deployment engineer (edge AI)
Builds
Production-ready, low-latency inference pipelines for autonomous vehicle stacks
Domain
Autonomous driving / Edge AI / Model Optimization
Deliverable
production ML models
Required skills
Model quantization (PTQ, QAT), Mixed-precision inference, Model conversion/compilation (TensorRT, ONNX), Custom CUDA kernel development, C++ (14/17/20), Python, FlashAttention, KV-cache optimization, Speculative Decoding
Preferred skills
Multi-modal sensor fusion, BEV/3D Occupancy Networks, Distributed training pipelines, End-to-end autonomous driving paradigms
Technologies
TensorRT, PyTorch, CUDA, C++, Python, ONNX, FlashAttention, DeepSpeed, Megatron-LM, Ray
Responsibilities
Optimize large-scale models using quantization, pruning, and parameter-efficient fine-tuning; Architect and implement model conversion and compilation pipelines; Perform parity checking, accuracy recovery, and latency benchmarking; Develop custom ML OPs and TensorRT Plugins with efficient CUDA kernels; Write production-level, low latency, and memory-safe C++ and CUDA code for real-time inference
Seniority
Senior, hands-on IC