Staff Site Reliability Engineer
Core
Design and implement high-reliability, scalable platform infrastructure systems handling billions of events daily for an AI marketing platform.
Role type
Staff Site Reliability Engineer (Infrastructure)
Builds
Compute, persistence, networking, observability, and deployment systems for a global AI marketing platform.
Domain
AI/ML Marketing Platform Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Production engineering, backend engineering, SRE, DevOps, strategic vision, complex problem solving, coding in Golang/Python/Java/Typescript, SLIs/SLOs, incident management, cross-team collaboration, technical mentorship
Preferred skills
Experience in dynamic, reliability-focused production environments
Technologies
Golang, Python, Java, Typescript
Responsibilities
Design and implement systems for reliability, observability, traceability, and incident management; Lead cross-team strategic initiatives; Establish production standards and best practices; Champion reliability goals via SLIs/SLOs; Mentor team members and develop engineering leaders
Seniority
Staff, strategic IC with mentorship
