Site Reliability Engineer (SRE)
Core
Own production reliability for a real-time platform combining voice, desktop, intelligence, and AI, where uptime and latency are the product.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Real-time communication and AI agent platforms
Domain
Telecommunications, AI, and Real-time Systems
Deliverable
production ML models | product features | infrastructure
Required skills
SLO and error budget ownership, incident response and command, production scaling, observability (metrics/logs/traces), event-driven and real-time systems reliability, Go or Python, networking protocols (TCP/UDP/TLS/WebSocket/DNS), capacity modeling, chaos engineering, AI-aware reliability monitoring, master debugging
Preferred skills
RTP/SIP knowledge, multi-cloud literacy, multi-tenancy isolation experience
Technologies
Go, Python, GCP, NATS, WebSocket, Kubernetes (implied by SRE context), Prometheus/Grafana (implied by observability)
Responsibilities
Own SLOs and error budgets per tenant/service, lead incident response and blameless postmortems, manage production scaling and capacity planning, implement observability depth for event hops, execute on-call rotation with DevOps, perform chaos engineering and failure injection, partner with DevOps on deploy-safety and canary analysis, monitor AI model latency and drift
Seniority
Senior, hands-on IC
