Senior Site Reliability Engineer
Core
Bridge the gap between Development, Cloud Platform Engineering, and Product Owners to define SRE concepts, align service quality with business objectives, and ensure reliable digital offerings.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Digital offerings, cloud-based systems, and self-developed software
Domain
Industrial Quality Solutions, Medical Technology, Vision Care, Sports & Cine Optics
Deliverable
production ML models | product features | infrastructure
Required skills
SRE concepts (SLI, SLO, SLA), observability setup, CI/CD pipeline automation, incident management, system architecture design, disaster recovery planning, infrastructure as code (Terraform), container orchestration (Kubernetes), monitoring and logging tools
Preferred skills
Cultural change advocacy, staying updated with latest SRE trends
Technologies
Kubernetes, Prometheus, Grafana, Open Telemetry Collector, Terraform
Responsibilities
Define and measure service reliability using SLI/SLOs; enable rapid software deployment while maintaining SLAs; setup observability and alerting for self-developed software and cloud components; maintain documentation, playbooks, and runbooks for troubleshooting; monitor system performance and conduct post-incident reviews; create and maintain disaster recovery plans; advocate for SRE best practices
Seniority
Senior, hands-on IC