Senior Software Engineer, Site Reliability Engineering, Cloud Storage
Core
Build and run large-scale, massively distributed, fault-tolerant systems for Google Cloud services, ensuring reliability, uptime, and performance.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Google Cloud's internally critical and externally-visible systems
Domain
Cloud infrastructure / Distributed systems
Deliverable
production ML models | infrastructure
Required skills
software development, large-scale distributed system design, system troubleshooting, project leadership, automation, capacity planning, incident response, system monitoring
Preferred skills
Master's degree in Computer Science or Engineering
Technologies
(none explicitly listed)
Responsibilities
Engage in the whole lifecycle of services from inception to refinement; Support services pre-launch via system design consulting and capacity planning; Maintain live services by monitoring availability and latency; Scale systems sustainably through automation; Practice blameless postmortems
Seniority
Senior, hands-on IC with leadership responsibilities