Senior Site Reliability Engineer - Data Infrastructure
Core
Engineering resilience, scalability, and efficiency for core data services and AI infrastructure powering products.
Role type
Senior Site Reliability Engineer (Data Infrastructure)
Builds
Production data services, AI infrastructure, and automation tools
Domain
Data Infrastructure / Distributed Systems
Deliverable
production ML models
Required skills
Linux/Unix, networking (TCP/IP, DNS), distributed systems, programming/scripting (Go, Python, Bash), Kubernetes, incident response, SLO/SLA management, capacity planning, automation, AI orchestration
Preferred skills
MySQL, Redis, Kafka, Flink, data center operations, complex incident leadership
Technologies
Kubernetes, Redis, MySQL, Kafka, Flink, Go, Python, Bash
Responsibilities
Incident response and postmortems, SLO/SLA and error budget management, capacity and cost optimization, pragmatic automation and AI orchestration, operational excellence and change management, data center and AI infrastructure construction, cross-team influence and mentorship
Seniority
Senior, hands-on IC
