Staff Software Engineer - Compute
Core
Design and lead the development of a resilient, highly available GPU and CPU host instance lifecycle control plane for a massive-scale AI cloud infrastructure.
Role type
Staff Software Engineer (Compute Infrastructure)
Builds
Next-generation GPU/CPU host instance lifecycle and compute control plane
Domain
Cloud Infrastructure / AI Compute / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Distributed systems, Durable execution models, Semiconductor architecture, BIOS/Firmware integration, DPU utilization, Multi-tenant security, Technical leadership, Data center deployment
Preferred skills
Nvidia AI Factory architecture, Linux kernel internals, Device drivers, Virtualization (KVM/QEMU), Kernel bypass (SR-IOV/DPDK/SPDK), High-performance networking (InfiniBand/RoCE), NVMe-oF
Technologies
C/C++, Rust, Python, Go, CUDA, DOCA, InfiniBand, RoCE, NVMe-oF
Responsibilities
Design and implement a highly available GPU/CPU host and instance lifecycle control plane; Guide technical decisions on semiconductor architecture, BIOS/Firmware, and DPU utilization; Lead design of compute platform multi-tenant security model; Provide technical leadership and mentorship for senior engineers; Collaborate with product and data center organizations to translate requirements into scalable capabilities; Set engineering standards and lead design reviews for mission-critical cloud software.
Seniority
Staff, hands-on IC with strategic leadership