Senior Software Engineer, SRE / Observability Tooling
Core
Building SRE and observability tooling, automation, and standards to enable engineering teams to monitor and maintain the health of cloud-native conversational AI infrastructure.
Role type
Senior IC Site Reliability Engineer (Observability Tooling)
Builds
Internal observability platform, automation frameworks, and standards for dashboards, alerts, and incident response.
Domain
Cloud-native infrastructure, Conversational AI, Banking Technology
Deliverable
production ML models | infrastructure
Required skills
AWS, Kubernetes (EKS), Observability platforms (DataDog, Prometheus), SRE principles (SLOs, error budgets), Python or Go, CI/CD pipelines (ArgoCD, Github Actions), Terraform, Distributed systems troubleshooting
Preferred skills
None stated
Technologies
AWS, Kubernetes, Istio, EFK, Amazon Aurora, RabbitMQ, Amazon RDS, Amazon ElastiCache, DataDog, Github Actions, ArgoCD, Jenkins, Helm, Terraform, Python, Go
Responsibilities
Develop standards and automation for dashboards, alerts, and monitors as code; Partner with development teams to establish production and operational readiness; Build tooling for defining and reporting on SLOs and SLIs; Develop automation to reduce manual toil in observability workflows; Build and improve incident response tooling and workflows.
Seniority
Senior, hands-on IC