Site Reliability Engineer, Infrastructure - ThousandEyes
Core
Design and operate large-scale, highly available distributed systems to process growing volumes of telemetry data for the ThousandEyes Digital Experience Assurance platform.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
The ThousandEyes platform processing hundreds of millions of messages per hour across cloud, internet, and enterprise networks.
Domain
Network Telemetry & Digital Experience Assurance
Deliverable
production ML models | infrastructure
Required skills
Python or Go, AWS, Kubernetes, Terraform, GNU/Linux administration, AI tooling for automation, incident management, capacity planning, root-cause analysis
Preferred skills
Distributed system design, SLO/SLA management, disaster-recovery testing, automation for release safety
Technologies
AWS, Kubernetes, Terraform, Python, Go, GNU/Linux
Responsibilities
Design and operate distributed systems for telemetry data processing, use AI tooling to automate solutions and reduce operational expense, evaluate scalability and security of production services, troubleshoot complex infrastructure issues, participate in on-call rotation and incident management, collaborate with development teams to meet SLOs and SLAs
Seniority
Senior, hands-on IC
