Staff Software Engineer, Platform Infrastructure
Core
Establish Site Reliability Engineering as a first-class discipline and build a self-service internal Service Foundation platform to improve reliability and observability for 15+ engineering teams.
Role type
Staff Software Engineer (Platform Infrastructure / SRE)
Builds
Internal Service Foundation platform, SRE patterns, Terraform modules, Go APIs, CLIs, and GCP-native managed services.
Domain
Cloud Infrastructure (GCP) / Site Reliability Engineering / Platform Engineering
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Site Reliability Engineering (SRE) principles, Cloud infrastructure (GCP), Observability (Prometheus, OpenTelemetry, Grafana), Infrastructure as Code (Terraform), Security mindset, Technical leadership, Platform engineering philosophy
Preferred skills
Go proficiency, Kubernetes operational experience, Security and networking expertise, Template and paved-path mindset
Technologies
Go, GCP (Cloud Run, GKE, Cloud SQL, Secret Manager), Terraform, Prometheus, OpenTelemetry, Grafana, Kubernetes
Responsibilities
Define and drive adoption of SLO/SLI frameworks and incident response practices; Own and evolve observability patterns; Design and deliver service templates and infrastructure patterns; Migrate to GCP-native managed services; Apply security-first thinking to networking and identity; Mentor engineers in reliability work; Operate in a full DevOps model; Participate in on-call rotation
Seniority
Staff, technical leadership & mentorship