CareerPlanSign in

Senior Software Engineer

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2026-09-01 → 2026-09-25

Core

DRI for supercomputing clusters and GPU compute/interconnect fabric operations, ensuring GPU availability, service reliability, and AI training stability.

Role type

Senior IC infrastructure engineer (HPC/AI systems)

Builds

Large-scale AI infrastructure, supercomputing clusters, GPU interconnect fabrics

Domain

High-performance computing (HPC), Artificial Intelligence (AI), Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

Cross-stack debugging, Linux systems administration, Hardware/Firmware/Driver/Software stack reasoning, Automation design, Telemetry implementation, Incident triage, Root cause analysis, Playbook development, Systemic failure pattern identification

Preferred skills

Deep knowledge of PCIe subsystems, GPU interactions, High availability design

Technologies

C, C++, C#, Java, JavaScript, Python, Linux

Responsibilities

Lead incident triage, mitigation, recovery, and root cause analysis for compute and fabric-related production issues; Perform deep, cross-stack debugging spanning hardware provisioning, GPU interconnect fabric, PCIe subsystems, and GPU interactions; Design and leverage automation, telemetry, and diagnostic tooling to improve issue detection, observability, and mean time to mitigation; Identify systemic failure patterns and develop technical guidance, troubleshooting procedures, and escalation frameworks

Seniority

Senior, hands-on IC

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.