Senior Software Engineer - Reliability Engineering (Remote)
Core
Ensure the resilience, performance, and security of the enterprise Cloud Platform by engineering reliability into platforms through automation, incident management, and destructive testing.
Role type
Senior Software Reliability Engineer (SRE)
Builds
Highly available, paved-path cloud platform solutions for product teams
Domain
Cloud Infrastructure & Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Cloud infrastructure automation, incident management, root cause analysis, SLO definition, destructive testing, container orchestration, scripting/programming, observability tooling, security frameworks
Preferred skills
AI tooling for incident response, ITIL processes, microservice architecture, performance testing, business impact reporting
Technologies
Terraform, Ansible, Google Cloud Platform, Kubernetes, GKE, Prometheus, Grafana, OpenTelemetry, BASH, Python, Golang, Typescript, Java, YAML, JSON, HCL, Wiz
Responsibilities
Develop, test, and deploy software; lead incident triage and root cause analysis; drive no-repeat resolutions via blameless postmortems; design and execute destructive and failure-scenario tests; mentor junior engineers on modern frameworks; manage live production incidents and report business impact
Seniority
Senior, hands-on IC