Software Engineer
Core
Senior Software Engineer guiding model serving, runtime tuning, accelerator optimization, and release evaluation for Cisco's Foundational Model Service.
Role type
Senior IC machine-learning infrastructure engineer (model serving)
Builds
High-availability platform serving small, large, and embedding models to Cisco engineering teams
Domain
AI infrastructure, LLM/SLM serving, GPU clusters
Deliverable
production ML models
Required skills
LLM/SLM/embedding model internals, vLLM/NVIDIA NIM/Triton runtimes, model evaluation and benchmarking, quantization and distillation, distributed GPU training, Kubernetes GPU workloads, Python or Go, CI/CD and IaC, observability tooling, mentoring
Preferred skills
AMD ROCm, KV cache aware routing, evaluation frameworks (lm-eval/DeepEval), API gateways (Envoy/NGINX), GitOps/Helm, capacity planning, open source inference contributions, CKA certification, PyTorch/Hugging Face fine-tuning, LoRA/QLoRA/PEFT, MLflow/W&B
Technologies
Kubernetes, vLLM, NVIDIA NIM, Triton, Prometheus, Grafana, Splunk, Terraform, Ansible, PyTorch, Hugging Face, Git, GitHub Actions
Responsibilities
Contribute to model serving direction and roadmap; guide quantization and accelerator optimization; develop platform services, APIs, and operational tooling; evolve evaluation and benchmarking frameworks; define model promotion criteria; evaluate fine-tuning and distillation effects; shape routing and capacity behavior; improve model registry and release workflows; develop observability; monitor production and lead postmortems; coordinate across teams; apply AI to platform operations; lead features from design to completion; write clean code and review for quality; mentor engineers and run design reviews; create technical designs and documentation
Seniority
Senior, hands-on IC with mentorship