Ingénieur fiabilité des infrastructures
Core
Maintaining, optimizing, and ensuring reliability and performance of critical SaaS infrastructure on AWS and Kubernetes with a focus on automation, observability, and continuous improvement.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Production SaaS platforms on AWS and Kubernetes for healthcare supply chain clients
Domain
Cloud Infrastructure / Healthcare Supply Chain
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident management, post-incident analysis (RCA), SLO/SLI definition, observability tooling, infrastructure as code (IaC), CI/CD pipelines, system design consultation, technical documentation, cross-team coordination
Preferred skills
None stated
Technologies
AWS, Kubernetes, Datadog, Terraform, GitLab CI/CD
Responsibilities
Collaborate with engineering teams on pre-launch system design and capacity planning; Identify weaknesses and lead initiatives to simplify and strengthen the platform; Monitor availability, latency, and system health; Optimize observability by defining SLO/SLI and creating actionable dashboards; Develop automation tools, IaC frameworks, and CI/CD pipelines to reduce manual intervention; Implement sustainable incident management and lead post-incident reviews (RCA); Act as incident commander during incidents to coordinate response and ensure rapid restoration
Seniority
Senior, hands-on IC