AI Senior Staff Systems Engineer
Core
Lead the end-to-end lifecycle of AI infrastructure, including architecting GPU clusters, deploying LLMs, and managing agentic AI workflows.
Role type
Senior Staff AI Systems Engineer
Builds
High-performance GPU clusters, production-grade agentic AI workflows, and scalable AI tech stacks
Domain
AI Infrastructure / High-Performance Computing
Deliverable
infrastructure
Required skills
NVIDIA GPU architecture (CUDA, cuDNN), public cloud AI services (Azure OpenAI, GCP), Docker/Kubernetes, Python/Bash/Perl scripting, Linux system administration, LLM deployment and optimization (quantization, distillation), job schedulers (LSF, Slurm)
Preferred skills
AI job profiling and tuning, macOS/AppleSilicon system administration
Technologies
NVIDIA GPUs, CUDA, cuDNN, Azure OpenAI, Google Cloud Platform, Docker, Kubernetes, PyTorch, TensorFlow, vLLM, TGI, TensorRT-LLM, LSF, Slurm, Python, Bash, Perl, Linux (RHEL), LDAP, Active Directory
Responsibilities
Design and implement next-generation AI infrastructure and technical strategy; manage secure access and billing for public cloud AI services; configure and optimize GPU server clusters; architect and deploy scalable AI tech stacks; lead deployment and optimization of Large Language Models; build production-grade agentic AI workflows; develop automation scripts and monitoring solutions; mentor engineers on AI systems best practices
Seniority
Senior Staff, hands-on individual contributor with strategic leadership