Lead Site Reliability Engineer
Core
Building scalable, resilient, and observable cloud-native systems for Tricentis SaaS products globally.
Role type
Lead Site Reliability Engineer (hands-on IC with leadership)
Builds
Cloud-native infrastructure supporting multi-region, multi-tenant deployments
Domain
SaaS / Public Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
SRE/DevOps leadership, public cloud operations (Azure), observability (SLOs/SLIs), infrastructure-as-code, container orchestration, CI/CD, incident management, technical mentoring
Preferred skills
GitOps, self-service platforms, chaos engineering, architectural decision making
Technologies
Azure, AWS, Terraform, GitHub Actions, Kubernetes, DataDog, Prometheus, Grafana, incident.io, Jira
Responsibilities
Lead cross-cutting initiatives for platform scalability and cost efficiency; Architect and implement cloud-native infrastructure; Improve observability strategy and alerting standards; Coach and mentor engineers; Own post-incident analysis; Influence product reliability from design to production; Establish deployment and incident response standards; Lead incident response for high-impact outages; Guide adoption of modern tooling and practices; Represent SRE in leadership forums
Seniority
Senior, hands-on IC with team leadership