AI Infrastructure Operations Engineer
Core
Design, build, and operate large-scale GPU and accelerated-computing infrastructure for AI training, inference, simulation, and high-performance compute workloads.
Role type
Senior IC AI Infrastructure Operations Engineer
Builds
Resilient, high-performance compute environments across cloud, on-premises, and hybrid deployments
Domain
AI Infrastructure / High-Performance Computing / Cloud Operations
Deliverable
production ML models | infrastructure
Required skills
GPU cluster management, Kubernetes orchestration, infrastructure automation, Terraform, Ansible, Python scripting, Bash scripting, capacity planning, incident management, performance benchmarking, distributed compute troubleshooting
Preferred skills
AgenticOps practices, REST API development, NVIDIA platform tools (Base Command Manager, NGC, NCCL, CUDA-X), LLM deployment on AI Cloud platforms, reliability engineering, change management
Technologies
Kubernetes, Slurm, Run:ai, Terraform, Ansible, Python, Bash, NVIDIA Base Command Manager, NGC, NCCL, CUDA-X
Responsibilities
Design and implement accelerated-computing infrastructure solutions; Deploy and operate GPU-based clusters; Integrate infrastructure platforms with enterprise systems; Build reusable automation tools and workflows; Establish repeatable operational processes; Perform and automate benchmarking and validation; Provide technical guidance and optimization for GPU clusters
Seniority
Senior, hands-on IC