Senior AI Infrastructure Engineer, Physical Infrastructure
Core
Lead the vision and execution of large-scale GPU compute infrastructure, ensuring high availability, fault tolerance, and automated resilience for ML training and inference clusters.
Role type
Senior hands-on IC AI Infrastructure Engineer (Physical & Logical)
Builds
Self-healing GPU clusters, high-performance interconnect fabrics, and automated deployment tooling for model training and inference.
Domain
Defense technology, AI/ML infrastructure, HPC
Deliverable
production ML models | infrastructure
Required skills
GPU compute at scale (H200/B200/B300), high-performance interconnects (NVLink, InfiniBand, RoCE, Spectrum-X), high-performance parallel storage (VAST, DDN, Weka), Kubernetes, infrastructure as code, automated deployment pipelines, observability and fault triage, physical datacenter operations (rack/stack/cable).
Preferred skills
NVIDIA NVL72 rack scale systems, LLM token serving/inference infrastructure, network fabric tuning (congestion control, adaptive routing, QoS), GPU/network observability tooling (DCGM, fabric telemetry).
Technologies
H200, B200, B300, NVL72, NVLink, InfiniBand, RoCE, Spectrum-X, VAST, DDN, Weka, Lustre, Kubernetes, Run:AI, Ray, DCGM.
Responsibilities
Rack, stack, cable, and bring up GPU compute systems including physical topology, power, cooling, and firmware management. Build and tune interconnect fabrics connecting hundreds of GPUs. Integrate high-performance parallel storage for distributed training. Automate cluster deployment and configuration end-to-end. Operate and extend Kubernetes/Run:AI environments for scheduling and isolation. Own fleet health monitoring and rapid triage of hardware/network faults. Onboard engineers and act as escalation point for infrastructure bottlenecks. Partner with product teams to translate compute needs into platform capabilities.
Seniority
Senior, hands-on IC