Senior Site Reliability Engineer, DGX Cloud
Core
Maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide.
Role type
Senior Site Reliability Engineer (AI Infrastructure)
Builds
Fully managed AI platform on major cloud providers
Domain
AI / High-Performance Computing / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes administration, containerization, microservices architecture, infrastructure automation (Terraform, Ansible, Chef, Puppet), high-level programming (Python, Go), Linux operating systems, networking fundamentals (TCP/IP), cloud security standards, SRE principles (SLOs, SLIs, error budgets), observability stacks (monitoring, logging, tracing)
Preferred skills
GPU-accelerated cluster operation with KubeVirt, generative-AI techniques for operational toil reduction, workflow orchestration (Temporal, Cadence, Airflow, Argo Workflows, Step Functions), production AI inference workload troubleshooting (vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL), GPU performance analysis
Technologies
Kubernetes, Terraform, Ansible, Chef, Puppet, Python, Go, Linux, OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, KubeVirt, Temporal, Cadence, Airflow, Argo Workflows, Step Functions, vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL
Responsibilities
Build, implement, and support operational and reliability aspects of large-scale Kubernetes clusters; Define SLOs/SLIs and monitor error budgets; Support services pre-launch via system creation consulting and capacity management; Maintain live services by measuring availability, latency, and system health; Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds; Scale systems sustainably through automation; Lead triage and root-cause analysis of high-severity incidents; Participate in on-call rotation for production services
Seniority
Senior, hands-on IC