CareerPlanSign in

Senior AI Inference Engineer - Model Optimization & Deployment

Foster City, CA💼 Full-time🗓 2026-04-11 → 2026-09-26

Core

Optimizing and deploying large-scale multi-modal foundation models (LLMs, VLMs, sensor fusion) for real-time inference on power- and thermal-constrained vehicle SOCs.

Role type

Senior IC model optimization and deployment engineer (edge AI)

Builds

Production-ready, low-latency inference pipelines for autonomous vehicle stacks

Domain

Autonomous driving / Edge AI / Model Optimization

Deliverable

production ML models

Required skills

Model quantization (PTQ, QAT), Mixed-precision inference, Model conversion/compilation (TensorRT, ONNX), Custom CUDA kernel development, C++ (14/17/20), Python, FlashAttention, KV-cache optimization, Speculative Decoding

Preferred skills

Multi-modal sensor fusion, BEV/3D Occupancy Networks, Distributed training pipelines, End-to-end autonomous driving paradigms

Technologies

TensorRT, PyTorch, CUDA, C++, Python, ONNX, FlashAttention, DeepSpeed, Megatron-LM, Ray

Responsibilities

Optimize large-scale models using quantization, pruning, and parameter-efficient fine-tuning; Architect and implement model conversion and compilation pipelines; Perform parity checking, accuracy recovery, and latency benchmarking; Develop custom ML OPs and TensorRT Plugins with efficient CUDA kernels; Write production-level, low latency, and memory-safe C++ and CUDA code for real-time inference

Seniority

Senior, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.