Staff AI/ML Infrastructure Engineer
Core
Designing and maintaining high-performance GPU and bare metal infrastructure to support large-scale AI and machine learning workloads.
Role type
Staff AI/ML Infrastructure Engineer
Builds
Scalable GPU clusters and automated testing frameworks for GPU-enabled platforms
Domain
Hardware infrastructure, GPU computing, and AI/ML systems
Deliverable
infrastructure
Required skills
Bare metal infrastructure, hardware automation, NVIDIA/AMD GPU platforms, high-performance networking (RoCE, InfiniBand), BIOS/BMC/firmware management, Linux systems administration, device drivers, Python, Bash, infrastructure automation, GPU driver optimization, multi-cluster environment management
Preferred skills
Machine learning software stack exposure, GPU-based workload optimization, hardware vendor partnership experience
Technologies
Bash, Firmware, InfiniBand, Linux, Python, PCIe
Responsibilities
Design and maintain GPU and bare metal infrastructure across containerized and physical environments, build scalable GPU clusters, ensure dependable high-performance provisioning, develop automated testing frameworks, implement infrastructure solutions for AI/ML workloads, benchmark and troubleshoot GPU performance at scale, partner with hardware vendors on drivers and firmware, resolve hardware and performance issues, optimize rail and cluster performance, provide technical leadership and mentor engineers