Software Engineer
Core
Owns end-to-end reliability for Azure Storage hardware in on-prem lab environments, including on-site work, and leads incident response for DPU-related issues.
Role type
Senior Site Reliability Engineer (DPU/Hardware)
Builds
Automation for provisioning, configuration, validation, canarying, rollback, patching, and recovery of DPU-enabled Azure Storage systems in large internal labs.
Domain
Cloud Infrastructure / Hardware / DPU / On-prem Lab
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Root cause analysis, Infrastructure automation, Infrastructure-as-code, Deployment pipelines, Observability, Alerting, Telemetry, Health modeling, SLO/SLI definition, Incident response, Cross-functional collaboration
Preferred skills
Fungible DPU technology, SmartNIC/DPU platforms, Powershell, Bash, Python
Technologies
Azure Monitor, DPU, SmartNIC, BIOS, Firmware, Networking, Operating Systems
Responsibilities
Lead live-site incident response and mitigation for hardware-, firmware-, or DPU-related issues; Build automation for provisioning, configuration, validation, canarying, rollback, patching, and recovery; Create and maintain operational runbooks, diagnostics, telemetry, and health models; Drive improvements in observability and alerting by extending Azure Monitor and internal systems; Partner with silicon, firmware, BIOS, networking, and OS teams to enable and validate DPU hardware; Supports platform bring-up, validation, and operational readiness activities for emerging hardware and infrastructure technologies.
Seniority
Senior, hands-on IC