Senior Site Reliability Engineer, Production Engineer - ThousandEyes
Core
Design and manage large-scale, highly available distributed systems in the cloud to enhance the reliability, performance, and security of the ThousandEyes Digital Experience Assurance platform.
Role type
Senior Site Reliability Engineer (Production Engineering)
Builds
Cloud-native services, scalable operations tooling, and automated production operations for a multi-region SaaS platform.
Domain
Cloud Infrastructure / Observability / SaaS
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Python or Go, Unix/Linux systems, Site Reliability principles, Cloud-native architecture, Incident response, Distributed systems, SLOs
Preferred skills
Kubernetes ecosystem, AWS cloud providers, Large-scale enterprise platform operations
Technologies
Kubernetes, Service Mesh, Prometheus, OpenTelemetry, ArgoCD, AWS
Responsibilities
Collaborate with software engineers to optimize architecture for availability and latency; Design and implement scalable operations tooling; Design, deploy, and maintain AWS cloud-native services; Participate in 24x7 incident response and on-call rotation; Automate production operations and deployment processes; Manage rapidly growing infrastructure handling substantial daily data volumes.
Seniority
Senior, hands-on IC