Senior Site Reliability Engineer
Core
Architect high-performing platforms and ensure reliability for an AI-powered learning platform serving millions of users globally.
Role type
Senior Site Reliability Engineer II (hands-on IC)
Builds
Scalable, resilient cloud infrastructure and observability tools for a SaaS e-learning platform
Domain
SaaS / E-learning / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
distributed systems, cloud architecture (AWS), container orchestration, Infrastructure as Code, incident management, observability, automation, mentoring
Preferred skills
self-healing infrastructure patterns, zero-downtime deployment pipelines, multi-tenant SaaS scaling
Technologies
AWS, container orchestration, Infrastructure as Code
Responsibilities
Define reliability objectives (SLIs/SLOs) and build long-term reliability roadmaps; lead cross-service incident responses; drive adoption of reliability patterns and deployment safety; build advanced observability tools; influence product design for reliability balance; mentor senior engineers and assist in recruiting
Seniority
Senior, hands-on IC