Site Reliability Engineering (SRE) Architect - CRL - Germany
Core
Design and implement reliability and scalability strategy for production systems, ensuring resilience, performance, and high availability for global clients.
Role type
Senior IC SRE Architect
Builds
Robust, scalable, and fault-tolerant infrastructure and application services on public cloud platforms
Domain
Cloud-native infrastructure, distributed systems, and enterprise reliability engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Cloud platform expertise (AWS/GCP/Azure), Kubernetes orchestration, Infrastructure as Code (Terraform), Observability strategy design, Distributed systems architecture, Python or Go programming
Preferred skills
Multi-cloud experience, Service mesh technologies (Istio/Linkerd), DevSecOps practices, Large-scale technology transformation leadership
Technologies
AWS, GCP, Azure, Kubernetes, Terraform, Ansible, Prometheus, Grafana, OpenTelemetry, Jaeger, ELK Stack, Datadog, New Relic, Istio, Linkerd
Responsibilities
Design long-term vision for system reliability and performance; Establish and govern SLOs, SLIs, and error budgets; Architect comprehensive observability strategies for logging, metrics, and tracing; Lead automation and IaC strategy using Terraform and Ansible; Design resilience patterns and chaos engineering experiments; Mentor SREs and developers on reliability best practices; Analyze major incidents to drive architectural improvements
Seniority
Senior, hands-on IC with strategic influence