CareerPlanSign in

Principal Supercomputing Operations Software Engineer

United States, Multiple Locations, Multiple Locations💼 Full-time💰 $142,800–$142,800🗓 2026-07-16 → 2026-09-25

Core

Technical authority and strategic owner for InfiniBand and GPU interconnect fabric operations across flagship AI supercomputing environments, ensuring reliability, performance, and SLA compliance for frontier AI training and large-scale simulations.

Role type

Principal Supercomputing Operations Software Engineer (IC5)

Builds

Hyperscale GPU clusters and interconnect fabrics for Azure AI/HPC

Domain

Cloud AI / High Performance Computing / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

C, C++, C#, Java, JavaScript, Python, Linux systems debugging, hardware/firmware/driver stack reasoning, large-scale distributed systems operations, incident command and triage, automation and telemetry architecture, failure modeling, playbook authoring

Preferred skills

Master's degree in Computer Science, 10+ years experience in HPC/AI infrastructure, deep expertise in InfiniBand and Subnet Manager

Technologies

InfiniBand, Subnet Manager, PCIe, GPUs, Linux, C, C++, C#, Java, JavaScript, Python

Responsibilities

Serve as DRI for interconnect fabric operations ensuring GPU availability and training stability; Lead and orchestrate complex high-severity fabric incidents end-to-end; Perform deep multi-layer systems debugging across hardware, firmware, drivers, and OS; Drive operational excellence by defining reliability models and authoring TSGs/playbooks; Architect automation, diagnostics, and telemetry to improve observability and debuggability

Seniority

Principal, hands-on IC with strategic ownership and mentorship

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.