AI Engineer, Inference
Core
Build and improve self-hosted AI inference services, establishing the engineering foundation for model serving in the organization's AI-factory environment to provide reliable, secure, scalable, and high-performance endpoints.
Role type
Senior AI Engineer (Inference)
Builds
Self-hosted model serving infrastructure, inference endpoints, and deployment templates for internal products and future Inference-as-a-service offerings.
Domain
AI Infrastructure / High-Performance Computing
Deliverable
production ML models
Required skills
Model serving optimization, distributed inference, Kubernetes, inference frameworks (TensorRT-LLM, vLLM, Triton), performance benchmarking, quantization, GPU memory management, observability
Preferred skills
Agentic workflows, speculative decoding, topology-aware placement
Technologies
TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, CUDA, cuDNN, NCCL, Kubernetes
Responsibilities
Build and operate self-hosted AI inference services; Define and implement standard model-onboarding workflows; Provision and manage secure, scalable inference endpoints; Develop reusable deployment templates, APIs, and SDKs; Optimize model-serving performance using quantization, compilation, and batching; Design distributed inference configurations; Work with Kubernetes and scheduler teams to define resource profiles; Build benchmarking and qualification workflows; Establish automated performance-regression testing; Build operational observability for inference services.
Seniority
Senior, hands-on IC