Staff Site Reliability Engineer
Core
Designing and engineering cloud infrastructure and operations solutions for a global IoT platform serving millions of homeowners and businesses.
Role type
Staff Site Reliability Engineer (Infrastructure & Platform)
Builds
Highly available global cloud platform (myQ) and PaaS implementations
Domain
IoT, Cloud Infrastructure, SRE
Deliverable
production ML models | infrastructure
Required skills
Cloud platform expertise (AWS), Infrastructure as Code (Terraform, ArgoCD), Kubernetes, Linux/Windows system administration, Observability (Prometheus, Grafana, Datadog, New Relic), Software development, Network architecture design, Disaster recovery, Security architecture
Preferred skills
Cloud engineering certifications, IoT industry experience, Team leadership, Cloud cost optimization, Advanced network security (VPC, firewall management)
Technologies
AWS, Terraform, ArgoCD, Kubernetes, Prometheus, Grafana, Datadog, New Relic, PowerShell, Python, Go, Bash, Active Directory
Responsibilities
Design and engineer infrastructure environments to solve challenges and create future opportunities; Lead initiatives from inception to implementation and optimization; Establish relationships with business members and technical staff to execute the infrastructure roadmap; Provide forward-thinking technology solutions with an emphasis on automation; Evolve PaaS implementations working with vendor solutions; Participate in monitoring, troubleshooting, and RCA for security and infrastructure events; Provide leadership, assistance, and technical guidance to the SRE Team; Collaborate with security teams on architecture design and tools; Maintain professional and technical knowledge through continuous learning.
Seniority
Staff, hands-on IC with indirect supervisory responsibilities