Technical Program Manager - Compute Systems Engineering
Core
Coordinate complex technical programs to sustain, scale, and improve operational and software development processes for a full-stack AI cloud platform's compute systems stack.
Role type
Senior Technical Program Manager (Compute Systems Engineering)
Builds
Nebius HPC and GPU Cloud deployments, including virtualization platforms, networking, and NVIDIA GPU clusters.
Domain
Cloud infrastructure, High-Performance Computing (HPC), Accelerated Computing, AI infrastructure
Deliverable
production ML models | infrastructure
Required skills
Program coordination, dependency management, risk identification, cross-functional leadership, process improvement, technical communication, stakeholder management, hands-on problem solving
Preferred skills
Cloud computing, HPC, accelerated computing, large-scale infrastructure, device drivers, Linux kernel, virtualization technologies, firmware lifecycle management
Technologies
Linux kernel, KVM, QEMU, NVIDIA GPUs, HPC networking
Responsibilities
Adapt systems stack to support latest hardware platforms, develop performance improvements and customer-facing capabilities, build deployment automation and validation processes, ensure observability across the stack, support customers with complex technical issues, participate in major hardware bring-ups and system validations
Seniority
Senior, hands-on IC