Software Engineer Manager - Platform Reliability Engineering (Remote)
Core
Lead a dedicated team of engineers to ensure the resilience, performance, and security of the enterprise cloud foundation through automation, rigorous change management, and systematic incident management.
Role type
Senior IC manager (Platform Reliability Engineering)
Builds
Enterprise cloud foundation services, reliability engineering practices, and paved-path solutions for product teams
Domain
Cloud infrastructure, Site Reliability Engineering (SRE), Platform Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Object-oriented programming (Java), team leadership and mentoring, SRE/DevOps team management, incident response and post-mortems, SLO/SLI definition and enforcement, public cloud ecosystem expertise, microservice architecture troubleshooting, observability stacks, container orchestration (Kubernetes/GKE), Infrastructure as Code (Terraform), chaos/resiliency testing, cloud networking and Zero Trust principles
Preferred skills
Cultural change driving, automated remediation SLAs, dynamic-scaling strategies, vendor/partner management, enterprise tooling implementation
Technologies
Java, Google Cloud Platform (GCP), Kubernetes, GKE, Terraform
Responsibilities
Collaborate with product teams to create secure, reliable, scalable software solutions; write custom code/scripts to automate infrastructure and monitoring; provide leadership, mentoring, and coaching to software engineers; establish and enforce Service Level Objectives (SLOs); evaluate new technologies for enterprise adoption; lead review board sessions to drive consistency; field technical questions and act as an escalation point; conduct annual and mid-year performance reviews
Seniority
Senior, hands-on IC with management responsibilities