Software Engineering Manager II, Site Reliability Engineering
Core
Lead a team of Software/Systems Engineers to ensure uptime, availability, and performance of large-scale distributed systems and Google's services.
Role type
Senior IC manager (Site Reliability Engineering)
Builds
Large-scale, fault-tolerant distributed systems and infrastructure for Google's services
Domain
Internet services, distributed systems, cloud infrastructure
Deliverable
production ML models | infrastructure
Required skills
distributed systems design, system troubleshooting, automation, code debugging, performance optimization, team leadership, project management
Preferred skills
expertise in computing, storage, or networking, large-scale system design, algorithmic complexity analysis
Technologies
distributed systems, data centers, Google platforms
Responsibilities
Lead a team of Software/Systems Engineers on projects for users and be responsible for uptime; Own end-to-end availability and performance of key services and build automation to prevent problem recurrence; Manage on-call rotations across continents using a follow-the-sun model; Design, write and deliver software to improve the availability, scalability, latency and efficiency of Google's services
Seniority
Senior, hands-on IC with management responsibilities