Senior AI Compute Engineer - NVIS
Core
Deploying, managing, and validating large-scale AI Compute/HPC infrastructure in Linux-based environments for customers.
Role type
Senior IC infrastructure engineer (AI compute systems)
Builds
Large-scale AI Compute systems and HPC clusters
Domain
Data center infrastructure, high-performance computing, AI hardware
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, cluster management, scripting (Bash, Python, Ansible), network routing and tuning, performance optimization, benchmarking tools (HPL, NCCL, MLPerf), job schedulers (SLURM, LSF, UGE), Kubernetes
Preferred skills
InfiniBand, GPU-focused hardware/software, MPI, storage technologies (Lustre, GPFS), OEM GPU platforms, Base Command Manager (BCM)
Responsibilities
Deploy and manage AI Compute/HPC infrastructure, act as domain expert during customer planning and implementation, provide feedback to internal teams on bugs and improvements, perform knowledge transfers for customer rollouts
Seniority
Senior, hands-on IC