Senior Site Reliability Engineer
Core
Design and operate resilient cloud-native data platforms for Crunchyroll's consumer-facing experiences, focusing on reliability, scalability, performance, and security.
Role type
Staff Site Reliability Engineer (SRE)
Builds
Cloud-native data platforms powering streaming, video, and interactive experiences for 100M+ users
Domain
Streaming media / Data & Insights / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes, GCP, Infrastructure as Code (Terraform), Linux systems administration, distributed systems, Go/Python/Java/Shell, Prometheus/Grafana/OpenTelemetry/Datadog, SLIs/SLOs, vulnerability management, penetration testing support, cloud/container security, OWASP Top 10, IAM, secrets management
Preferred skills
SecOps best practices, disaster recovery strategy, capacity planning, self-healing mechanisms, SSDLC
Technologies
Kubernetes, GCP, Terraform, Prometheus, Grafana, OpenTelemetry, Datadog, Go, Python, Java, Shell
Responsibilities
Define and improve reliability via SLIs/SLOs/error budgets; manage incidents and postmortems; build observability (logging/tracing/alerting); develop automation and self-service capabilities; optimize cloud-native infrastructure scalability; implement IaC and deployment automation; plan capacity and performance; validate disaster recovery strategies; integrate security controls and remediate vulnerabilities; support penetration testing and secure environment operations.
Seniority
Staff, hands-on IC with service leadership
