Systems Development Engineer, AWS Incident Response (AIR)
Core
Lead real-time response to critical customer-impacting events on AWS infrastructure while building automation tools to prevent recurrence.
Role type
Senior Systems Development Engineer (Incident Response)
Builds
Incident detection, triage, and mitigation automation tools; dashboards for real-time event visibility.
Domain
Cloud Infrastructure / Systems Engineering / Incident Management
Deliverable
production ML models | product features | infrastructure
Required skills
Systems engineering fundamentals, Networking, Storage systems, Operating systems, Modern programming languages (C++, C#, Java, Python, Golang, PowerShell, Ruby), Design patterns, Reliability engineering, Scaling systems
Preferred skills
Large enterprise technical customer escalations, Cross-functional project delivery, Root cause analysis, Infrastructure automation
Technologies
C++, C#, Java, Python, Golang, PowerShell, Ruby
Responsibilities
Lead incident calls and coordinate resolver teams during large-scale events, Design and build incident response automation tools, Author COEs and event deep-dive documents, Identify recurring platform issues and own projects to eliminate operational problems, Collaborate globally to expand incident response capabilities
Seniority
Senior, hands-on IC