CareerPlanSign in

Solutions Architect, Inference Deployments

US, CA, Santa Clara💼 Full-time💰 $152,000–$152,000🗓 2026-04-15 → 2026-09-25

Core

Design and deploy scalable AI inference solutions using NVIDIA GPU technology and Kubernetes for enterprise customers.

Role type

Senior Solutions Architect (AI Inference)

Builds

Production-grade generative AI inference pipelines and disaggregated inference systems

Domain

AI/ML Infrastructure, GPU Computing, Cloud Native

Deliverable

production ML models

Required skills

Distributed systems architecture, Kubernetes orchestration, GPU resource management, LLM optimization, low-latency networking, technical leadership

Preferred skills

NVIDIA Dynamo, Triton Inference Server, TensorRT-LLM, vLLM, SGLang, Transformer neural networks, quantization, speculative decoding, open-source contributions

Technologies

NVIDIA Dynamo, Kubernetes, TensorRT-LLM, vLLM, SGLang, NVIDIA GPU Operator, NIM Operator, MIG, RDMA, UCX

Responsibilities

Build inference pipelines with NVIDIA Dynamo, orchestrate disaggregated inference using Kubernetes, accelerate pipelines with TensorRT-LLM/vLLM/SGLang, provide mentorship and technical leadership for deployments

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.