Senior Sr Staff AI Infrastructure
Core
Architect and optimize vehicle-side AI model deployment, focusing on low-latency execution of LLMs and foundational models within embedded constraints.
Role type
Senior Staff AI Infrastructure Engineer (Autonomous Driving)
Builds
Service-oriented deployment environments for LLMs and foundational models supporting offline simulation, automated annotation, and model validation.
Domain
Autonomous Driving / Embedded AI Systems / High-Performance Computing
Deliverable
production ML models
Required skills
C++, Python, CUDA programming, OpenMP, low-level system profiling, TensorRT, ONNX Runtime, vLLM, SGLang, TensorRT-LLM, GPU architecture knowledge (NVIDIA Hopper/Thor), memory bandwidth management, quantization techniques, kernel fusion, graph compilation
Preferred skills
NVIDIA Thor optimization, serving foundation models with vLLM/SGLang/TGI/LightLLM, PyTorch, real-time AI workload deployment in robotics or edge devices
Technologies
CUDA, OpenMP, TensorRT, ONNX Runtime, vLLM, SGLang, TensorRT-LLM, PyTorch, NVIDIA Hopper, NVIDIA Thor
Responsibilities
Own deployment, optimization, and resource scheduling of vehicle-side AI models; Lead vehicle-side system stability initiatives including root-cause analysis; Architect and scale service-oriented deployment environments; Evaluate and integrate optimization toolchains and execution engines; Establish profiling and telemetry frameworks using CUDA tools; Collaborate with Perception, Cloud Infrastructure, and Safety teams on algorithm iteration and deployment
Seniority
Senior Staff, hands-on IC with strategic architecture responsibilities