Senior Software Engineer
Core
Develop and maintain tools and systems for service reliability, monitoring, alerting, and incident management at scale.
Role type
Senior IC Site Reliability Engineer (Production SRE)
Builds
Incident management platform, reliability tooling, and automation for engineering teams
Domain
Cloud infrastructure, distributed systems, SRE
Required skills
Java, Python, or Go; distributed systems; cloud platforms (AWS/GCP); containerization (Docker/Kubernetes); automated testing; CI/CD
Preferred skills
Incident command leadership; mentoring engineers; cross-functional crisis coordination
Technologies
AWS, Google Cloud Platform, Docker, Kubernetes, Java, Python, Go
Responsibilities
Design and implement reliability tools and monitoring systems; lead high-severity incidents as Incident Commander; mentor engineers on incident handling; collaborate with infrastructure teams to solve operational challenges; participate in post-mortems to address systemic issues.
Seniority
Senior, hands-on IC