Technical Operations & Site Reliability Engineer, Customer Systems
Core
Design and maintain automation solutions for large-scale, globally distributed customer systems to ensure reliability, availability, and performance.
Role type
Senior IC Technical Operations & Site Reliability Engineer
Builds
Automated monitoring, incident response, and operational workflows for business-critical global applications
Domain
Technology / Distributed Systems / Customer Experience
Deliverable
production ML models
Required skills
Java/JEE, REST, Swift/Objective C, Python, Go, Bash, Linux kernel management, networking protocols (HTTP, DNS, TCP/IP, ICMP), log analysis, incident response, AI/LLM model training and optimization, database schema design
Preferred skills
Microservices architecture, messaging brokers, versioning strategies, 24x7 global operations experience, system health monitoring strategy
Technologies
Hubble, ExtraHop, Splunk, Linux, AI & LLM models
Responsibilities
Manage large-scale production outages and lead incident response; Develop tools to automate repetitive operational tasks; Plan and execute system health monitoring and communication; Partner with teams to improve reliability and processes; Create and maintain technical documentation and training materials
Seniority
Senior, hands-on IC