Site Reliability Engineer
Core
Designing and operating distributed systems primitives, auto-scaling infrastructure, and observability for a platform running AI agents and background workflows at massive scale.
Role type
Senior / Staff Site Reliability Engineer
Builds
Production infrastructure for AI agents and background workflows
Domain
Developer infrastructure, AI agents, distributed systems
Deliverable
production ML models | infrastructure
Required skills
Distributed systems design, observability (OpenTelemetry, Prometheus), TypeScript, Go, self-managed Kubernetes, Terraform, security hardening, Postgres, Redis
Preferred skills
AWS, performance tuning, multi-tenant isolation
Technologies
OpenTelemetry, Prometheus, Kubernetes, Terraform, TypeScript, Go, Postgres, Redis, AWS
Responsibilities
Owning observability across the platform, designing and operating distributed systems primitives, architecting and tuning auto-scaling infrastructure, hardening multi-tenant sandbox isolation, owning Terraform and IaC, running on-call practice
Seniority
Senior / Staff, hands-on IC
