Staff Site Reliability Engineer
Core
Founding member of a new SRE team ensuring reliability, performance, and availability for Develocity, a toolchain observability platform serving paying customers and open-source projects.
Role type
Staff Site Reliability Engineer (Founding Member)
Builds
Production SaaS instances of Develocity, artifact registries, and supporting infrastructure on AWS.
Domain
DevOps / Cloud Infrastructure / SRE
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes, AWS (EKS, RDS, S3, EC2), observability tools (Prometheus, Grafana), Infrastructure as Code (Terraform), scripting (Python, Bash), incident management, SLOs and error budgets, automation, technical leadership, cross-team influence
Preferred skills
Founding/early SRE experience, JVM languages (Java, Kotlin), customer-facing incident communications
Technologies
Kubernetes, AWS, Prometheus, Grafana, Terraform, Python, Bash, Develocity
Responsibilities
Operate and maintain Develocity instances and supporting services in production; Define and evolve SRE standards, practices, and operating models; Participate in on-call rotation as technical escalation point; Lead incident response and blameless retrospectives; Set reliability priorities using risk and SLOs; Identify systemic reliability risks and evolve SaaS operations; Lead architectural reviews for reliability and scalability; Drive automation across deployment, upgrades, and recovery; Build comprehensive observability; Own disaster recovery and business continuity; Partner with engineering leadership on reliability; Mentor and coach SREs; Communicate with customers during incidents; Optimize performance and operational costs.
Seniority
Staff, hands-on IC with leadership and mentorship responsibilities