Member of Technical Staff, Site Reliablity Engineer
Core
Building reliability infrastructure and platform services for a real-time voice AI platform that handles live phone calls.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Auto-remediation services, capacity forecasters, oncall tooling, and autoscaling logic for voice call workers.
Domain
Real-time voice AI / Telecommunications infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Incident command, SLOs and error budgets, capacity planning, load testing, Kubernetes production ops, backpressure and autoscaling patterns
Preferred skills
Go or TypeScript development, real-time latency-sensitive product experience, postmortem discipline
Technologies
Go, TypeScript, Bash, Chronosphere, Prometheus, Grafana, Datadog, OpenTelemetry, Kubernetes (EKS), KEDA
Responsibilities
Run incident command and manage oncall rotation, define and monitor SLOs for call completion, tune autoscaling and capacity planning, ship platform services for reliability, conduct load testing and postmortems
Seniority
Senior, hands-on IC