Staff Site Reliability Engineer
Core
Architect core reliability platforms and drive AI-driven reliability (AIOps) for a mission-critical SaaS platform serving 6,000+ enterprise customers.
Role type
Staff Site Reliability Engineer (hands-on IC with mentorship)
Builds
Reliability paved paths, self-service tooling, and AIOps automation for global engineering teams
Domain
Enterprise SaaS / Market Intelligence / AI
Deliverable
production ML models | infrastructure
Required skills
Python or Go, Kubernetes, AWS/GCP/Azure, TCP/IP/DNS/HTTP/S, Prometheus/Grafana/OTEL, incident command, full-stack troubleshooting
Preferred skills
AIOps strategy, continuous profiling, blameless postmortems
Technologies
Kubernetes, Prometheus, Grafana, OTEL, Datadog, ELK, Python, Go
Responsibilities
Architect reliability frameworks and self-service tooling; Lead AI-driven reliability and AIOps strategy; Champion SRE culture via design reviews and standards; Act as Incident Commander during critical events; Deliver end-to-end observability and profiling; Mentor engineers across SRE and product teams
Seniority
Staff, hands-on IC with strategic influence