Applied AI Engineer, Inference
Core
Build and optimize high-performance model serving capabilities for real production workloads, focusing on latency, throughput, and quality.
Role type
Applied AI Engineer (Inference Performance)
Builds
High-performance inference stack for AI models
Domain
AI Infrastructure / LLM Inference Systems
Deliverable
production ML models
Required skills
Python, empirical benchmarking, LLM inference systems (vLLM, SGLang, TensorRT-LLM), latency/throughput optimization, profiling, quantization, speculative decoding
Preferred skills
GPU hardware optimization, Nsight Systems/PyTorch profilers, production trace analysis, evaluation frameworks for coding/reasoning
Technologies
vLLM, SGLang, TensorRT-LLM, Nsight Systems, PyTorch
Responsibilities
Build and maintain benchmarking workflows for latency, throughput, and cost; Profile model-serving behavior to identify bottlenecks; Drive targeted optimizations for customer workloads; Design and run experiments on inference techniques; Partner with platform engineers to productionize improvements; Produce technical writeups for configuration and deployment decisions
Seniority
Mid-Senior, hands-on IC