Principal Supercomputing Operations Software Engineer
Core
Technical authority and strategic owner for InfiniBand and GPU interconnect fabric operations across flagship AI supercomputing environments, ensuring reliability, performance, and SLA compliance for frontier AI training and large-scale simulations.
Role type
Principal Supercomputing Operations Software Engineer (IC5)
Builds
Hyperscale GPU clusters and interconnect fabrics for Azure AI/HPC
Domain
Cloud AI / High Performance Computing / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
C, C++, C#, Java, JavaScript, Python, Linux systems debugging, hardware/firmware/driver stack reasoning, large-scale distributed systems operations, incident command and triage, automation and telemetry architecture, failure modeling, playbook authoring
Preferred skills
Master's degree in Computer Science, 10+ years experience in HPC/AI infrastructure, deep expertise in InfiniBand and Subnet Manager
Technologies
InfiniBand, Subnet Manager, PCIe, GPUs, Linux, C, C++, C#, Java, JavaScript, Python
Responsibilities
Serve as DRI for interconnect fabric operations ensuring GPU availability and training stability; Lead and orchestrate complex high-severity fabric incidents end-to-end; Perform deep multi-layer systems debugging across hardware, firmware, drivers, and OS; Drive operational excellence by defining reliability models and authoring TSGs/playbooks; Architect automation, diagnostics, and telemetry to improve observability and debuggability
Seniority
Principal, hands-on IC with strategic ownership and mentorship
