Staff Site Reliability Engineer
Core
Staff SRE ensuring observability, scalability, and operability for a global online IDE and AI-powered app builder serving millions of developers.
Role type
Staff Site Reliability Engineer (IC)
Builds
WebContainers, Bolt.new (AI-powered app builder), and the underlying infrastructure for millions of developers.
Domain
Developer tools / Cloud Infrastructure / AI-powered IDE
Deliverable
production ML models | infrastructure
Required skills
Multi-cloud fluency (AWS, GCP, Azure), Infrastructure as Code (Terraform), TypeScript, Ruby on Rails, SRE/Production Engineering, Technical Leadership, Strategic Execution, Systems Thinking, Data-Driven Leadership
Preferred skills
SRE practice maturity at growth-stage companies, embedded SRE experience, chaos/resilience testing design
Technologies
AWS, GCP, Azure, Terraform, TypeScript, Ruby on Rails
Responsibilities
Embed with product and platform teams early in project lifecycle to design for reliability; Define production-readiness standards and operational acceptance criteria; Establish SLIs, SLOs, and error budgets; Build frameworks and tooling across multi-cloud environments; Lead incident management and blameless postmortems; Represent the company with cloud providers and in the reliability community.
Seniority
Staff, high-influence IC with strategic execution
