CareerPlanSign in

Perception Deployment Engineer - Model Deployment & Optimization

Foster City, CA💼 Full-time🗓 2026-09-30 → 2026-10-01

Core

Deploying and optimizing large-scale multi-modal foundation models (sensor fusion, LLMs, VLMs) for real-time execution on vehicle SOCs.

Role type

Senior IC perception deployment engineer (model optimization & edge deployment)

Builds

Production-ready, low-latency inference code and optimized model binaries for autonomous vehicle stacks

Domain

Autonomous driving, computer vision, edge AI

Deliverable

production ML models

Required skills

C++ (14/17/20), CUDA, model quantization (PTQ, QAT), TensorRT, mixed-precision inference, custom ML OPs, PyTorch, ONNX, latency benchmarking

Preferred skills

FlashAttention, KV-cache optimization, BEV, 3D Occupancy Networks, VLM/VLA models, TensorRT-LLM

Responsibilities

Design and develop production-level C++ and CUDA code for real-time perception algorithms; Optimize large-scale models using quantization and mixed-precision frameworks; Architect and implement model conversion and compilation pipelines using TensorRT; Perform parity checking, accuracy recovery, and latency benchmarking; Develop and optimize custom ML OPs and TensorRT Plugins with efficient CUDA kernels

Sourced via lever · Listed on CareerPlan, which tracks 902,000+ jobs from 20+ sources.