Senior Site Reliability Engineer
Core
Strengthen infrastructure and enhance ability to deploy, monitor, and scale systems effectively and reliably for over 50 Product Engineering teams.
Role type
Senior Site Reliability Engineer (IC)
Builds
Multi-tenant SaaS platform infrastructure, cloud-native systems, and internal platform tools
Domain
Cloud infrastructure, SaaS, Digital Employee Experience
Deliverable
production ML models | product features | infrastructure
Required skills
Cloud services (AWS, GCP, Azure), Kubernetes, Infrastructure-as-Code (Terraform), CI/CD pipelines, Monitoring solutions, Linux systems, Networking, Scripting (Python, Go, Bash), Zero-downtime deployment strategies
Preferred skills
Service mesh (Istio), Chaos engineering, Compliance standards (SOC 2, ISO 27001, HIPAA, FedRAMP), Cost optimization
Technologies
AWS, Kubernetes, Terraform, Docker, Helm, Jenkins, GitHub Actions, GitLab CI, FluxCD, Crossplane, Datadog, Istio, S3, EBS
Responsibilities
Implement and manage cloud-native systems; Operate and enhance Kubernetes clusters and service meshes; Design and maintain infrastructure for multi-tenant SaaS; Define and maintain SLOs, SLAs, and error budgets; Develop infrastructure-as-code; Build internal platform tools and automation; Monitor infrastructure and applications; Participate in on-call rotation and act as Incident Commander; Drive incident response processes; Diagnose and resolve complex issues; Work with software engineers to embed reliability principles; Automate runbooks and alerting; Support automated testing and deployment strategies; Contribute to security best practices and cost optimization.
Seniority
Senior, hands-on IC