Senior Software Engineer II, Developer Experience / Operational Excellence
Core
Designing and building automated reliability, self-healing, and observability systems for a globally distributed engineering organization to ensure production stability and resilience.
Role type
Senior IC platform engineer (operational excellence)
Builds
Automated rollbacks, deploy safeguards, fault mitigation tooling, incident management systems, and AI-driven operational tooling for internal engineering teams.
Domain
Cloud infrastructure, DevOps, SRE, Observability, AI-driven automation
Deliverable
production ML models | infrastructure
Required skills
Architecting monitoring frameworks and SLO platforms, designing automated response workflows, building observability infrastructure, implementing AI-driven automation in SDLC, mentoring engineers, cloud platform expertise, Go/Python for infrastructure code
Preferred skills
Incident management tooling (PagerDuty, Incident.io), Infrastructure as Code (Terraform)
Technologies
Datadog, New Relic, Grafana, AWS, GCP, Terraform, Go, Python
Responsibilities
Design and build automated reliability and self-healing systems; Own and improve incident management tooling and on-call health; Develop and evolve observability infrastructure; Contribute to AI-driven operational tooling; Drive incident prevention by identifying systemic patterns; Partner with product engineering teams to diagnose reliability gaps; Define and champion operational excellence best practices
Seniority
Senior, hands-on IC with mentorship responsibilities