SRE- Availability Engineer
Core
Senior technical leader responsible for leading incident response, triage, and restoration for critical production systems during outages and degradations.
Role type
Senior Site Reliability Engineer (Technical Duty Officer)
Builds
Incident response capabilities, operational resilience, and system availability standards
Domain
E-commerce, large-scale internet platforms
Deliverable
client delivery
Required skills
Incident leadership, operational judgment, cross-functional coordination, risk assessment, automation development, system troubleshooting
Preferred skills
Incident commander experience, e-commerce platform experience, executive communication, monitoring improvement, telemetry analysis
Technologies
Go, Python, Java, Node.js, Docker, Kubernetes
Responsibilities
Lead technical response during high-priority incidents, maintain awareness of critical service health, serve as senior technical escalation point, partner to improve availability through automation and tooling, drive operational excellence via documentation and post-incident follow-through
Seniority
Senior, hands-on IC