Senior Software Engineer
Core
Investigate service incidents, analyze root causes, and implement systemic improvements to ensure network reliability and resilience.
Role type
Senior Software Engineer (Network Reliability & Incident Response)
Builds
Mitigation steps, Incident Learning Reviews (ILRs), postmortems, and long-term corrective actions for network services.
Domain
Enterprise Networking, High-Performance Computing, Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Incident investigation and root cause analysis, Network troubleshooting and diagnostics, High-performance packet processing, eBPF and XDP programming, MsQuic optimization, Rebootless update mechanisms, Cross-functional collaboration, Technical documentation
Preferred skills
AI application for efficiency, System design understanding, Mentoring peers
Technologies
C, C++, C#, Java, Python, TCP/IP, UDP, DNS, DHCP, HTTP/HTTPS, QUIC, Wireshark, procmon, Hyper-V, SDN, NMR, eBPF, XDP, MsQuic
Responsibilities
Identify root causes of incidents through data analysis and develop mitigation steps. Prepare and present Incident Learning Reviews (ILRs) and postmortems. Partner with Engineering teams to verify fix effectiveness. Analyze incident trends to drive proactive remediation. Participate in on-call rotations for assigned service queues. Collaborate on complex code-level or architecture-level networking issues.
Seniority
Senior, hands-on IC