Site Reliability Engineer
Core
Lead scaling operational resilience, stability, observability, and debugging workflows for an AI workforce orchestration platform serving global enterprises.
Role type
Senior Site Reliability Engineer
Builds
Internal tooling for on-call teams, reliability workflows, and observability pipelines
Domain
AI infrastructure, distributed systems, enterprise operations
Deliverable
production ML models | infrastructure
Required skills
Go, Kubernetes, debugging production systems, observability tools (Grafana, Prometheus, Sentry), problem-solving under pressure
Preferred skills
distributed systems at scale, internal tooling development, CI/CD pipelines, infra-as-code, custom metrics and traces
Technologies
Go, Kubernetes, Grafana, Prometheus, Sentry
Responsibilities
Own system stability and uptime, design tools to reduce incident load, debug complex real-time failures, shift operations from reactive to proactive
Seniority
Senior, hands-on IC