Senior AI Infrastructure Engineer - DGX Cloud
Core
Design, build, and maintain large-scale production systems for AI training and inferencing platforms using cloud infrastructure and distributed systems.
Role type
Senior IC infrastructure engineer (cloud & distributed systems)
Builds
Internal tooling and platforms for large-scale AI training and inferencing
Domain
High-performance computing, cloud infrastructure, AI systems
Deliverable
production ML models | infrastructure
Required skills
distributed systems architecture, infrastructure automation, capacity management, system design, incident response, performance analysis, Linux, networking, storage, containers, IaC, Python, Go, C/C++, Java, Kubernetes, Terraform
Preferred skills
experience with Slurm, large-scale private/public cloud operations, automation of routine tasks
Technologies
Kubernetes, OpenStack, Terraform, Slurm, Linux
Responsibilities
Design, build, deploy, and run internal tooling for AI platforms; conduct performance characterization on multi-GPU clusters; manage capacity and system health; support services through design consulting and launch reviews; maintain live services via monitoring; scale systems sustainably; practice blameless postmortems; participate in on-call rotation
Seniority
Senior, hands-on IC