Solutions Architect, Inference Deployments
Core
Design and deploy scalable AI inference solutions using NVIDIA GPU technology and Kubernetes for enterprise customers.
Role type
Senior Solutions Architect (AI Inference)
Builds
Production-grade generative AI inference pipelines and disaggregated inference systems
Domain
AI/ML Infrastructure, GPU Computing, Cloud Native
Deliverable
production ML models
Required skills
Distributed systems architecture, Kubernetes orchestration, GPU resource management, LLM optimization, low-latency networking, technical leadership
Preferred skills
NVIDIA Dynamo, Triton Inference Server, TensorRT-LLM, vLLM, SGLang, Transformer neural networks, quantization, speculative decoding, open-source contributions
Technologies
NVIDIA Dynamo, Kubernetes, TensorRT-LLM, vLLM, SGLang, NVIDIA GPU Operator, NIM Operator, MIG, RDMA, UCX
Responsibilities
Build inference pipelines with NVIDIA Dynamo, orchestrate disaggregated inference using Kubernetes, accelerate pipelines with TensorRT-LLM/vLLM/SGLang, provide mentorship and technical leadership for deployments
Seniority
Senior, hands-on IC