CareerPlanSign in

NIM Solution Architect

China, Shanghai💼 Full-time🗓 2026-09-21 → 2026-09-25

Core

Design, implement, and optimize NVIDIA Inference Microservices (NIM) for enterprise LLM, VLM, and generative AI workloads across on-prem, cloud, and hybrid environments.

Role type

Senior Solution Architect (AI Inference & Model Optimization)

Builds

Containerized NIM microservices, optimized inference pipelines, and technical demos for enterprise clients.

Domain

AI Infrastructure / Generative AI / High-Performance Computing

Deliverable

production ML models

Required skills

NIM microservices deployment, LLM/VLM inference optimization, distributed GPU training, PyTorch, transformer architecture tuning, rollout sampling strategies, model evaluation

Preferred skills

RL rollout frameworks (SLIME, Nemo-RL), programmatic verification/simulators, enterprise AI deployment, agent systems

Technologies

NVIDIA NIM, PyTorch, Python, GPU clusters

Responsibilities

Package and serve open-source and proprietary models via standardized APIs, tune NIM models for high-volume inference, deliver technical projects and client support, collaborate on expanding the AI solutions portfolio

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.