Staff Site Reliability Engineer
Core
Senior technical leader driving reliability strategy, leading high-risk technical initiatives, and setting engineering standards for complex platform components.
Role type
Staff Site Reliability Engineer (Strategic IC)
Builds
High-availability cloud-native platforms, CI/CD pipelines, and AI/ML-powered solutions
Domain
Cloud Infrastructure, Site Reliability Engineering, Platform Engineering
Deliverable
production ML models | infrastructure
Required skills
Cloud Architecture, Kubernetes, CI/CD pipeline design, Linux internals, Database architecture, Networking, Infrastructure-as-Code, Scripting/Automation, AI/ML platform integration
Preferred skills
LLM orchestration, Vector databases, Model serving infrastructure, AI observability
Technologies
GCP, Kubernetes, Datadog, Prometheus, Grafana, PagerDuty, Terraform, Pulumi, Python, Bash, Go, PostgreSQL, MySQL, Bigtable, Firestore, Redis
Responsibilities
Lead research, testing, and implementation for new systems and tooling; Perform complex capacity planning, load testing, and security improvements; Model calm incident response and lead high-risk maintenance events; Elevate team standards through new tooling and processes; Represent the team during transitions or coverage gaps
Seniority
Staff, strategic leadership with hands-on execution
