Senior Site Reliability Engineer (x/f/m)
Core
Ensure Doctolib's platform remains reliable, scalable, and resilient at a European scale across infrastructure, observability, and cross-cutting reliability initiatives for 170+ applications.
Role type
Senior Site Reliability Engineer (IC)
Builds
Cloud-native infrastructure automation, reliability components, and observability systems for a healthcare platform.
Domain
Healthcare technology / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, Infrastructure as Code, Cloud platforms (AWS/GCP/Azure), Incident response, SLO/SLI/Error Budgets, GitOps (ArgoCD/Helm), Scripting/Programming (Python/Go/Ruby)
Preferred skills
Experience in regulated environments (healthcare/fintech), Reliability enablement practices
Technologies
Kubernetes, Terraform, AWS, GCP, Prometheus, OpenTelemetry, Datadog, ArgoCD, Helm
Responsibilities
Build and maintain infrastructure automation and infrastructure-as-code at scale; Identify and lead large-scale cross-cutting reliability initiatives; Design, build, and improve infrastructure components for reliability and observability; Define and drive SLOs, error budgets, and alerting standards; Participate in on-call rotation and improve incident response; Partner with software engineering teams to embed reliability practices.
Seniority
Senior, hands-on IC