AI/ML Technical Leader - Language Model Inference & AI Ops
Core
Build and operate scalable AI systems for Intelligent Customer Experiences, moving LLM/SLM capabilities from prototype to production across cloud and on-prem environments.
Role type
Senior IC AI/ML Technical Leader (Inference & AI Ops)
Builds
Production-grade AI platforms, model-serving pipelines, and inference optimization services for Cisco's CX Incubation team.
Domain
Enterprise AI / Large Language Models / Cloud & On-Prem Infrastructure
Deliverable
production ML models
Required skills
Python, Java, C++, PyTorch, TensorFlow, GPU inference optimization, CI/CD for models, model observability, on-prem deployment packaging, speculative decoding, continuous batching, quantization (F8/INT4), multi-GPU parallelism, PEFT/LoRA, LLM evaluation
Preferred skills
vLLM, TensorRT-LLM, Triton, SGLang, llama.cpp, Nsight, PyTorch profiler, speculative/assisted decoding, paged/flash attention, KV-cache management, GPTQ/AWQ/SmoothQuant, tensor/pipeline/expert parallelism, disaggregated serving, ONNX Runtime, OpenVINO, MLC, K8s, model registry, experiment tracking
Technologies
vLLM, TensorRT-LLM, Triton, SGLang, llama.cpp, Nsight, PyTorch profiler, ONNX Runtime, OpenVINO, MLC, K8s
Responsibilities
Build robust model-serving and deployment pipelines with SLAs and rollback strategies; Optimize inference performance using speculative decoding, continuous batching, and quantization; Package and integrate on-prem inference stacks with secure configuration; Design scalable serving architectures including tensor/pipeline parallelism; Build automated CI/CD for models and prompts; Implement model and service observability including latency metrics and drift detection; Support training and fine-tuning workflows including data curation and experiment tracking.
Seniority
Senior, hands-on IC