CareerPlanGet AI match score →

Senior Inference Engineer, AIConfigurator for Dynamo

2 Locations💼 Full-time💰 $184,000–$184,000🗓 2026-06-12 → 2026-07-30

Core

Build and evolve AIConfigurator, a system that automatically discovers high-performance deployment configurations for large-scale LLM inference on NVIDIA platforms.

Role type

Senior Inference Engineer (LLM serving & optimization)

Builds

AIConfigurator optimization engine, production APIs/CLIs/SDKs, and configuration generation artifacts for GPU clusters.

Domain

AI Infrastructure / LLM Inference / GPU Systems

Deliverable

production ML models | product features | infrastructure

Required skills

Python, Rust, GPU computing, distributed systems, LLM inference concepts, performance modeling, benchmarking, software architecture

Preferred skills

TensorRT-LLM, vLLM, SGLang, Dynamo, Kubernetes, H100/H200/B200/GB200 experience, disaggregated serving, NCCL/NIXL/NVSHMEM, open-source contribution

Technologies

Python, Rust, Kubernetes, TensorRT-LLM, vLLM, SGLang, Dynamo, Triton Inference Server, NCCL, NIXL, NVSHMEM

Responsibilities

Build core optimization engine for LLM serving including configuration search and SLA-aware ranking; Develop production-quality APIs, CLIs, and SDKs for deployment configuration generation; Create systems to emit backend-specific artifacts for various serving platforms; Collaborate with runtime and platform teams to validate simulated results against actual deployment performance; Integrate performance databases and profiling data to improve model and hardware support; Drive software quality through architecture, schema development, testing, and automation; Convert complex inference concepts into dependable software abstractions

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗