Senior Site Reliability Engineer (AI)
Core
Own production services end-to-end, ensuring reliability, scalability, and operational excellence for a highly available and performant application platform.
Role type
Senior Site Reliability Engineer (AI)
Builds
The OneTrust AI-Ready Governance Platform™
Domain
Enterprise data governance and AI regulation
Deliverable
production ML models | infrastructure
Required skills
Incident response, SLI/SLO definition, error budget management, system reliability, cross-functional collaboration
Preferred skills
None stated
Technologies
None stated
Responsibilities
Own production services end-to-end including reliability, scalability, and operational excellence; Participate in on-call rotation and lead incident response; Collaborate with Engineering, Operations, and Product teams to design, deliver, and maintain the platform; Define, implement, and maintain SLIs and SLOs aligned with customer experience; Design and instrument SLIs such as latency, error rates, and availability across critical services; Manage and enforce error budgets to balance system reliability with product feature velocity.
Seniority
Senior, hands-on IC