Senior Site Reliability Engineer (f/m/d)
Core
Build and maintain a self-service runtime platform and developer tooling to enable 70+ engineers to write code, run workloads, and ship software reliably on GCP.
Role type
Senior Site Reliability Engineer (Platform Engineering)
Builds
Self-service runtime platform, developer portal, CI/CD pipelines, observability solutions, and disaster recovery capabilities.
Domain
Cloud Infrastructure (GCP), Kubernetes, Platform Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Backend/Infrastructure engineering, SRE/Platform engineering, GCP/AWS, Kubernetes, Terraform, Helm, TypeScript, Observability (Datadog), SLO/Error Budgets, Infrastructure as Code (IaC), GitOps, Distributed systems design
Preferred skills
None stated
Technologies
GCP, Kubernetes, Terraform, Helm, TypeScript, Datadog, MongoDB
Responsibilities
Build runtime platform as a self-service product, own developer portal and internal platform roadmap, ensure site reliability via observability and disaster recovery, define and operate SLOs and error budgets, drive infrastructure cost optimization, improve security posture, act as second line of defense for incidents
Seniority
Senior, hands-on IC