Site Reliability Engineer
Core
Building and maintaining scalable, reliable infrastructure and automation tools for a SaaS platform serving 60M+ locations globally.
Role type
Senior Site Reliability Engineer
Builds
Cloud-native infrastructure, CI/CD pipelines, and observability solutions for a smart home service delivery platform
Domain
Cloud infrastructure / Smart home / ISP services
Deliverable
production ML models | infrastructure
Required skills
Python, Go, AWS, GCP, Terraform, Salt, Kubernetes, Prometheus, Grafana, OpenSearch, IaC, observability, CI/CD design, cloud cost optimization
Preferred skills
large-scale distributed systems, networking, security best practices, SQL, NoSQL
Responsibilities
Implement and maintain scalable infrastructure using IaC; Develop observability solutions for high availability; Design and maintain CI/CD pipelines; Collaborate with development teams on operability; Participate in global on-call rotation; Drive operational toil reduction through automation; Optimize cloud resources for cost and performance
Seniority
Senior, hands-on IC