Site Reliability Engineer
Core
Design, build, and maintain large-scale, fault-tolerant distributed systems to ensure platform reliability and performance.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Large-scale, massively distributed, fault-tolerant systems for the Genesis platform
Domain
Cloud computing and distributed systems
Deliverable
production ML models | infrastructure
Required skills
Software engineering (Python, Go), distributed systems design, cloud platforms (Kubernetes, Cloud Functions), system optimization, automation, capacity planning, performance analysis
Preferred skills
Technical leadership, complex project management, cross-team collaboration
Technologies
Python, Go, Kubernetes, Cloud Functions
Responsibilities
Drive solutions for reliability, scalability, and efficiency challenges; optimize existing systems and eliminate toil through automation; ensure long-term service health via capacity planning and proactive incident prevention; guide technical decisions balancing system health with product priorities
Seniority
Senior, hands-on IC with technical leadership