Principal Software Engineer
Core
Design and build GPU-accelerated infrastructure for training and inference workloads, managing device scheduling, isolation, and sharing across bare metal, VMs, and containers.
Role type
Principal Software Engineer (GPU Infrastructure & Orchestration)
Builds
GPU device management systems, orchestration platforms (AKS, DRA, Kubernetes), virtualization stacks, and high-performance interconnects for AI workloads.
Domain
Cloud Infrastructure, AI/ML Platforms, GPU Computing, Virtualization
Deliverable
production ML models | infrastructure
Required skills
GPU device management, Kubernetes ecosystem, high-performance networking (RDMA/InfiniBand), virtualization and containerization, system design for large-scale fleets, observability and diagnostics, technical leadership, cross-team architectural alignment, debugging complex cross-layer systems.
Preferred skills
Experience with confidential compute, multitenant AI platforms, GPU virtualization/passthrough/partitioning technologies.
Technologies
AKS, Dynamic Resource Allocation (DRA), Kubernetes, C, C++, C#, Java, JavaScript, Python, RDMA, InfiniBand
Responsibilities
Design and build GPU-accelerated infrastructure for training and inference; develop systems for GPU device management, scheduling, isolation, and sharing; build and operate advanced orchestration and resource governance scenarios; evolve virtualization and container stacks for modern AI workloads; optimize performance and utilization across large GPU fleets; partner with networking and storage teams for high-performance interconnects; drive end-to-end platform features including observability and operational excellence; influence platform architecture and technical direction through design reviews.
Seniority
Principal, hands-on IC with strategic influence and mentorship