Director - Backend Engineering - AI Infra
Core
Lead the design and delivery of software-defined infrastructure (SDN orchestrators, fleet management, storage backends) to power global AI training and inference clusters.
Role type
Director, Backend Engineering (AI Infrastructure)
Builds
Automated GPU fleet management systems, high-performance storage backends, and SDN orchestrators for AI clusters
Domain
AI Infrastructure / High-Performance Computing (HPC)
Deliverable
infrastructure
Required skills
Linux internals, distributed systems, L2/L3 networking, Python, Go, C++, Terraform, Kubernetes, Ansible, parallel file systems (Lustre, Weka, VAST), GPU architecture
Preferred skills
Custom SDN controller development, NVIDIA/GPUDirect technologies, hyper-scale cloud environments (AWS, Azure, GCP, Meta)
Responsibilities
Lead design of SDN orchestrators for GPU networking; oversee backend services for GPU health and fault detection; drive backend logic for global traffic routing and load balancing; build backend interfaces for parallel file systems; direct strategy for AI object storage; act as final technical authority for AI Infra Architecture; champion Hardware-as-Code culture using Python, Ansible, and Terraform; lead multi-disciplinary org including Backend Developers and Infra Ops teams
Seniority
Director, strategic leadership with hands-on technical authority