CareerPlanGet AI match score →

AI Inference Performance Engineer

US, CA, Santa Clara💼 Full-time💰 $152,000–$152,000🗓 2026-03-09 → 2026-07-31

Core

Optimizing and benchmarking GenAI inference performance on NVIDIA accelerators for language, video, and speech workloads.

Role type

Senior IC GPU performance engineer (inference)

Builds

High-performance inference tools and benchmarks for large-scale LLMs, VLMs, and diffusion models

Domain

GPU deep learning / High-performance computing

Deliverable

production ML models

Required skills

Python, C++, PyTorch, JAX, LLM/VLM architecture, CUDA programming, distributed inference, roofline analysis, quantization, scheduling, memory management

Preferred skills

TensorRT-LLM, vLLM, SGLang, kernel development (CUTLASS, Triton), compiler/runtime optimization, scale-out orchestration (MPI, NCCL, K8S), team leadership

Technologies

TensorRT-LLM, SGLang, vLLM, CUDA, PyTorch, JAX, NCCL, K8S, MPI

Responsibilities

Drive end-to-end optimization pipelines for inference; Define and optimize next-generation AI workloads; Architect distributed inference from single-GPU to rack-scale clusters; Establish performance methodology via profiling and roofline analysis; Contribute to open-source inference frameworks; Lead technical execution and team performance

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗