Staff Site Reliability Engineer
Core
Building a high-scale infrastructure team responsible for owning environments with thousands of nodes, unifying complex infrastructure stacks, and driving the transformation from monoliths to scalable microservices.
Role type
Staff Site Reliability Engineer (Infrastructure)
Builds
Scalable platform/API architectures and internal tools supporting developer experience
Domain
Cloud infrastructure, microservices, high-availability systems
Deliverable
production ML models | product features | infrastructure
Required skills
Cloud engineering (GCP/AWS), microservices architecture, observability, CI/CD, system design, technical leadership, mentorship
Preferred skills
Scaling infrastructure to thousands of nodes, leading architectural shifts from monoliths to microservices, deep knowledge of reliability engineering (9s), first-party infrastructure integration, experience at large-scale tech companies
Technologies
GCP, AWS, microservices, observability tools, CI/CD tools
Responsibilities
Execute transformation from monolith to scalable microservices, drive initiatives to improve reliability, architect systems enabling best practices by default, integrate diverse infrastructure components, design observability and CI/CD frameworks, collaborate cross-functionally, provide technical leadership and mentor a team of 6 engineers
Seniority
Staff, hands-on IC with strategic leadership