Senior Inference Engineer, AIConfigurator for Dynamo
Core
Build and evolve AIConfigurator, a system that automatically discovers high-performance deployment configurations for large-scale LLM inference on NVIDIA platforms.
Role type
Senior Inference Engineer (LLM serving & optimization)
Builds
AIConfigurator optimization engine, production APIs/CLIs/SDKs, and configuration generation artifacts for GPU clusters.
Domain
AI Infrastructure / LLM Inference / GPU Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Python, Rust, GPU computing, distributed systems, LLM inference concepts, performance modeling, benchmarking, software architecture
Preferred skills
TensorRT-LLM, vLLM, SGLang, Dynamo, Kubernetes, H100/H200/B200/GB200 experience, disaggregated serving, NCCL/NIXL/NVSHMEM, open-source contribution
Technologies
Python, Rust, Kubernetes, TensorRT-LLM, vLLM, SGLang, Dynamo, Triton Inference Server, NCCL, NIXL, NVSHMEM
Responsibilities
Build core optimization engine for LLM serving including configuration search and SLA-aware ranking; Develop production-quality APIs, CLIs, and SDKs for deployment configuration generation; Create systems to emit backend-specific artifacts for various serving platforms; Collaborate with runtime and platform teams to validate simulated results against actual deployment performance; Integrate performance databases and profiling data to improve model and hardware support; Drive software quality through architecture, schema development, testing, and automation; Convert complex inference concepts into dependable software abstractions
Seniority
Senior, hands-on IC