CareerPlanSign in

Software Engineer

United States, Washington, Redmond💼 Full-time🗓 2026-07-21 → 2026-09-26

Core

Owns end-to-end reliability for Azure Storage hardware in on-prem lab environments, including on-site work, and leads incident response for DPU-related issues.

Role type

Senior Site Reliability Engineer (DPU/Hardware)

Builds

Automation for provisioning, configuration, validation, canarying, rollback, patching, and recovery of DPU-enabled Azure Storage systems in large internal labs.

Domain

Cloud Infrastructure / Hardware / DPU / On-prem Lab

Deliverable

production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work

Required skills

Root cause analysis, Infrastructure automation, Infrastructure-as-code, Deployment pipelines, Observability, Alerting, Telemetry, Health modeling, SLO/SLI definition, Incident response, Cross-functional collaboration

Preferred skills

Fungible DPU technology, SmartNIC/DPU platforms, Powershell, Bash, Python

Technologies

Azure Monitor, DPU, SmartNIC, BIOS, Firmware, Networking, Operating Systems

Responsibilities

Lead live-site incident response and mitigation for hardware-, firmware-, or DPU-related issues; Build automation for provisioning, configuration, validation, canarying, rollback, patching, and recovery; Create and maintain operational runbooks, diagnostics, telemetry, and health models; Drive improvements in observability and alerting by extending Azure Monitor and internal systems; Partner with silicon, firmware, BIOS, networking, and OS teams to enable and validate DPU hardware; Supports platform bring-up, validation, and operational readiness activities for emerging hardware and infrastructure technologies.

Seniority

Senior, hands-on IC

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.