Staff Software Engineer, Site Reliability Engineering, Networking
Core
Design, analyze, and troubleshoot large-scale distributed systems and networking infrastructure to ensure reliability, uptime, and performance for Google Cloud services.
Role type
Staff Software Engineer, Site Reliability Engineering (Networking)
Builds
Large-scale, fault-tolerant distributed systems and networking infrastructure for Google Cloud
Domain
Cloud computing, distributed systems, networking
Deliverable
production ML models | product features | infrastructure
Required skills
Software development (Go, Java, C, C++), networking (WAN, edge, routing protocols), distributed systems design, project leadership, automation, system debugging and optimization
Preferred skills
Large-scale distributed systems troubleshooting, code optimization, routine task automation
Technologies
Go, Java, C, C++, WAN, edge, routing protocols
Responsibilities
Engage in and improve the whole life-cycle of services from inception and design through to deployment, operation, and refinement; Support services before they go live through system design consulting, developing software platforms and frameworks, capacity planning, and launch reviews; Maintain services once they are live by measuring and monitoring availability, latency, and overall system health; Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity; Practice sustainable incident response and blameless postmortems.
