Staff Site Reliability Engineer
Core
Senior technical authority defining strategy for Observability, Alerting, and Platform Infrastructure, shaping engineering culture and bridging business goals with internet-scale execution.
Role type
Staff Site Reliability Engineer (Strategic IC)
Builds
Reliable, scalable, secure cloud platforms and distributed systems
Domain
Legal Tech / AI / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
distributed systems, cloud infrastructure, observability, reliability engineering, AIOps, Kubernetes, Python/Go/Bash, Infrastructure as Code, incident response, capacity planning, automation, mentoring
Preferred skills
regulated environment experience (FedRAMP, CJIS, HIPAA, SOC 2, PCI)
Technologies
Kubernetes, New Relic, Datadog
Responsibilities
Define technical strategy for Observability & Alerting and Platform Infrastructure; Lead evolution of reliable cloud platforms; Champion SLIs, SLOs, error budgets, and automation; Lead complex production incidents and drive permanent improvements; Build self-service platform capabilities to reduce toil; Mentor engineers and serve as trusted technical authority.
Seniority
Staff, strategic IC with mentorship